How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("feature-extraction", model="optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True)
# Load model directly
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True)
model = AutoModel.from_pretrained("optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True, device_map="auto")
Quick Links

Tiny Random Qwen3 DFlash b16

Tiny random DFlash draft model compatible with optimum-intel-internal-testing/tiny-random-qwen3.

  • Source remote-code template: z-lab/Qwen3-4B-DFlash-b16
  • Target hidden size: 32
  • Target hidden layers: 2
  • DFlash block size: 16
  • DFlash target layer ids: [0, 1]
Downloads last month
11,015
Safetensors
Model size
30.9k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support