Vishva007/d3-mini-W4A16-AutoRound

This repository hosts a high-precision W4A16 quantized version of vllm-sr/d3-mini using Intel AutoRound.

d3-mini is the 4B multimodal foundation decision model of Decision 3.0 (vLLM Semantic Router). It evaluates inputs (text, JSON, images, and video clips) against decision questions (Choice, Yes/No, Score) and outputs direct probabilities in a single forward pass without generating tokens.


Quantization Details

The quantization was performed using AutoRound with a high-accuracy, production-grade configuration. To preserve zero-shot vision and video grounding fidelity, the vision encoder was explicitly kept unquantized.

Tuning Configuration

Parameter Value Note
Bit Width W4A16 4-bit weights, 16-bit activations
Group Size 32 Fine-grained quantization grouping
Symmetric (sym) True Symmetric signed quantization
Iterations (iters) 1200 High convergence / production-grade optimization
Samples (nsamples) 512 Calibration sample size
Batch Size 4 Calibration batch size
Sequence Length (seqlen) 4096 Max sequence context during calibration
quant_nontext_module False Keeps Vision Encoder in FP16/BF16 to prevent visual drift
enable_torch_compile True Accelerated calibration via TorchDynamo

Installation

Ensure you have the required dependencies and compatible transformers installed:

pip install "transformers==5.17.0" torch torchvision pillow opencv-python-headless safetensors accelerate
pip install auto-round
# Optional for linear-attention speedups:
pip install flash-linear-attention

Quickstart & Usage

You can load and query the model directly via the standard system_one decision API.

1. Pure Text Routing

import json
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "Vishva007/d3-mini-W4A16-AutoRound", 
    trust_remote_code=True,
    device_map="auto"
)

result = model.system_one(
    state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "returns": "Refunds, replacements and damaged deliveries",
                "billing": "Payments, invoices and charges",
                "technical": "Product setup and faults"
            }
        },
        "receipt": {
            "type": "noul",
            "instructions": "Does the customer have a receipt?"
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is this request?",
            "criteria": ["Routine", "Soon", "Today"]
        }
    },
)

print(json.dumps(result["answers"], indent=2))

2. Multimodal Routing (Images & Video)

Because the vision tower is preserved in unquantized precision (quant_nontext_module: False), full visual resolution and feature fidelity are retained:

from huggingface_hub import hf_hub_download

# Download sample asset from original repo
receipt_img = hf_hub_download("vllm-sr/d3-mini", "assets/example-receipt.png")

result = model.system_one(
    state="The customer says the blender arrived cracked and attached the receipt.",
    images=[receipt_img],  # Accepts paths, URLs, PIL images, or base64 strings
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "returns": "Refunds, replacements and damaged deliveries",
                "billing": "Payments, invoices and charges"
            }
        },
        "on_receipt": {
            "type": "noul",
            "instructions": "Does the receipt list the blender?"
        }
    },
)

print(result["answers"])

Memory & Performance Advantages

  • VRAM Footprint: Reduces base language backbone weight footprint by roughly 3–3.5× compared to standard BF16.
  • Edge Deployment: Designed to run efficiently on low-memory edge accelerators, discrete workstation GPUs, or embedded hardware.
  • Visual Accuracy: Zero degradation in visual QA benchmarks (CV-Bench, Perception Test) due to unquantized vision tower caching.

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.

PyTorch 2.14

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.14 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.14-runpod d7lxsa4w9m Deploy to RunPod
PyTorch 2.14 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.14-runpod yk0y6j6rpg Deploy to RunPod
PyTorch 2.14 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.14-runpod gsp4gwx0nw Deploy to RunPod

PyTorch 2.13

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.13 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.13-runpod gmlupxnxfk Deploy to RunPod
PyTorch 2.13 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.13-runpod y3j8xvk4f4 Deploy to RunPod
PyTorch 2.13 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.13-runpod vigpissn5w Deploy to RunPod

PyTorch 2.12

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.12 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.12-runpod ctmz86zmf0 Deploy to RunPod
PyTorch 2.12 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.12-runpod qjko5yiwzi Deploy to RunPod
PyTorch 2.12 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.12-runpod ifg6xmye0f Deploy to RunPod

License & Attribution

  • Base Model: vllm-sr/d3-mini (Apache-2.0) by the vLLM Semantic Router Team.
  • License: Apache-2.0.
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/d3-mini-W4A16-AutoRound

Finetuned
Qwen/Qwen3.5-4B
Finetuned
vllm-sr/d3-mini
Quantized
(3)
this model

Collection including Vishva007/d3-mini-W4A16-AutoRound