Text Generation
Transformers
Safetensors
English
qwen3
Mixture of Experts
text-generation-inference
code
deepscale
math
conversational
Instructions to use prithivMLmods/Segue-Qwen3_DeepScaleR-Preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prithivMLmods/Segue-Qwen3_DeepScaleR-Preview with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="prithivMLmods/Segue-Qwen3_DeepScaleR-Preview") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("prithivMLmods/Segue-Qwen3_DeepScaleR-Preview") model = AutoModelForCausalLM.from_pretrained("prithivMLmods/Segue-Qwen3_DeepScaleR-Preview", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prithivMLmods/Segue-Qwen3_DeepScaleR-Preview with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prithivMLmods/Segue-Qwen3_DeepScaleR-Preview
- SGLang
How to use prithivMLmods/Segue-Qwen3_DeepScaleR-Preview with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use prithivMLmods/Segue-Qwen3_DeepScaleR-Preview with Docker Model Runner:
docker model run hf.co/prithivMLmods/Segue-Qwen3_DeepScaleR-Preview
File size: 3,901 Bytes
dd2512c b28f4cc 81f6ffb 585b225 caea4a1 585b225 ca2eefe 585b225 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 | ---
license: apache-2.0
datasets:
- agentica-org/DeepScaleR-Preview-Dataset
base_model:
- Qwen/Qwen3-4B
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- moe
- text-generation-inference
- code
- deepscale
- math
---

# Segue-Qwen3\_DeepScaleR-Preview
> Segue-Qwen3\_DeepScaleR-Preview is an experimental fine-tuned variant of the Qwen3-4B model architecture. It is trained on the DeepScaleR-Preview dataset—comprising high-quality mathematical reasoning problems—to achieve exceptional performance in symbolic, mathematical, and logical tasks with lightweight computational requirements.
## Key Features
1. Precision Reasoning with DeepScaleR-Preview Dataset
Fine-tuned on approximately 40,000 curated math problem-answer pairs sourced from:
* AIME (1984–2023)
* AMC (pre-2023)
* Omni-MATH
This enables superior symbolic manipulation and step-by-step logical deduction.
2. Lightweight Code Understanding
Capable of interpreting and generating correct code in Python, C++, and other logic-intensive languages with an emphasis on problem-solving and structured thought.
3. Structured Output Formatting
Outputs are designed to be well-formatted in Markdown, JSON, LaTeX, or tables—ideal for technical documentation, math notebooks, and data workflows.
4. Instruction-Following Accuracy
Strong multi-step instruction adherence, particularly for STEM domains. Ensures continuity, factual correctness, and process transparency in reasoning chains.
5. Multilingual Capabilities
Supports over 20 languages for mathematical and logical reasoning, technical instruction translation, and cross-lingual academic support.
6. Efficient 4B Architecture
Built on the Qwen3-4B base model to balance performance and scalability. Runs efficiently on mid-range GPUs while delivering high-accuracy inference.
## Quickstart with Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "prithivMLmods/Segue-Qwen3_DeepScaleR-Preview"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
prompt = "Solve for x: 5(x - 2) = 3x + 4, showing all steps clearly."
messages = [
{"role": "system", "content": "You are a precise mathematical assistant trained on DeepScaleR-Preview dataset."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=512
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
```
## Intended Use
* Step-by-step mathematical problem solving
* Symbolic computation and logic derivation
* Code generation and correction in technical environments
* Automated LaTeX/Markdown/JSON generation for education and documentation
* Academic tutoring and educational assistants
* Multilingual reasoning and translation of structured content
## Limitations
* Less suitable for open-domain conversation or creative writing
* Smaller context window compared to large-scale LLMs
* May be sensitive to token formatting in edge-case symbolic prompts
* Could underperform on intentionally adversarial logic inputs
## References
1. Qwen2.5 Technical Report – [https://arxiv.org/pdf/2412.15115](https://arxiv.org/pdf/2412.15115)
2. YaRN: Context Window Extension for LLMs – [https://arxiv.org/pdf/2309.00071](https://arxiv.org/pdf/2309.00071) |