Instructions to use Vishva007/d3-mini-W4A16-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Vishva007/d3-mini-W4A16-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="Vishva007/d3-mini-W4A16-AutoRound", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("Vishva007/d3-mini-W4A16-AutoRound", trust_remote_code=True) model = AutoModel.from_pretrained("Vishva007/d3-mini-W4A16-AutoRound", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Vishva007/d3-mini-W4A16-AutoRound
This repository hosts a high-precision W4A16 quantized version of vllm-sr/d3-mini using Intel AutoRound.
d3-mini is the 4B multimodal foundation decision model of Decision 3.0 (vLLM Semantic Router). It evaluates inputs (text, JSON, images, and video clips) against decision questions (Choice, Yes/No, Score) and outputs direct probabilities in a single forward pass without generating tokens.
Quantization Details
The quantization was performed using AutoRound with a high-accuracy, production-grade configuration. To preserve zero-shot vision and video grounding fidelity, the vision encoder was explicitly kept unquantized.
Tuning Configuration
| Parameter | Value | Note |
|---|---|---|
| Bit Width | W4A16 | 4-bit weights, 16-bit activations |
| Group Size | 32 |
Fine-grained quantization grouping |
Symmetric (sym) |
True |
Symmetric signed quantization |
Iterations (iters) |
1200 |
High convergence / production-grade optimization |
Samples (nsamples) |
512 |
Calibration sample size |
| Batch Size | 4 |
Calibration batch size |
Sequence Length (seqlen) |
4096 |
Max sequence context during calibration |
quant_nontext_module |
False |
Keeps Vision Encoder in FP16/BF16 to prevent visual drift |
enable_torch_compile |
True |
Accelerated calibration via TorchDynamo |
Installation
Ensure you have the required dependencies and compatible transformers installed:
pip install "transformers==5.17.0" torch torchvision pillow opencv-python-headless safetensors accelerate
pip install auto-round
# Optional for linear-attention speedups:
pip install flash-linear-attention
Quickstart & Usage
You can load and query the model directly via the standard system_one decision API.
1. Pure Text Routing
import json
from transformers import AutoModel
model = AutoModel.from_pretrained(
"Vishva007/d3-mini-W4A16-AutoRound",
trust_remote_code=True,
device_map="auto"
)
result = model.system_one(
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"
}
},
"receipt": {
"type": "noul",
"instructions": "Does the customer have a receipt?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": ["Routine", "Soon", "Today"]
}
},
)
print(json.dumps(result["answers"], indent=2))
2. Multimodal Routing (Images & Video)
Because the vision tower is preserved in unquantized precision (quant_nontext_module: False), full visual resolution and feature fidelity are retained:
from huggingface_hub import hf_hub_download
# Download sample asset from original repo
receipt_img = hf_hub_download("vllm-sr/d3-mini", "assets/example-receipt.png")
result = model.system_one(
state="The customer says the blender arrived cracked and attached the receipt.",
images=[receipt_img], # Accepts paths, URLs, PIL images, or base64 strings
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges"
}
},
"on_receipt": {
"type": "noul",
"instructions": "Does the receipt list the blender?"
}
},
)
print(result["answers"])
Memory & Performance Advantages
- VRAM Footprint: Reduces base language backbone weight footprint by roughly 3–3.5× compared to standard BF16.
- Edge Deployment: Designed to run efficiently on low-memory edge accelerators, discrete workstation GPUs, or embedded hardware.
- Visual Accuracy: Zero degradation in visual QA benchmarks (CV-Bench, Perception Test) due to unquantized vision tower caching.
🚀 Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
PyTorch 2.13
PyTorch 2.12
License & Attribution
- Base Model:
vllm-sr/d3-mini(Apache-2.0) by the vLLM Semantic Router Team. - License: Apache-2.0.
- Downloads last month
- -