RLMs (Reasoning Language Models)
updated
LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
Paper
• 2503.00735
• Published • 23
START: Self-taught Reasoner with Tools
Paper
• 2503.04625
• Published • 113
R1-Searcher: Incentivizing the Search Capability in LLMs via
Reinforcement Learning
Paper
• 2503.05592
• Published • 27
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with
Reinforcing Learning
Paper
• 2503.05379
• Published • 38
21B • Updated • 294
• 389
Viewer
• Updated • 269 • 1.1k
• 47
Text Generation
• 33B • Updated • 374k
• • 2.95k
Text Generation
• 8B • Updated • 295
• • 185
Text Generation
• 33B • Updated • 35
• • 157
Reinforcement Learning for Reasoning in Small LLMs: What Works and What
Doesn't
Paper
• 2503.16219
• Published • 52
predibase/Predibase-T2T-32B-RFT
33B • Updated • 6
• 20
agentica-org/DeepCoder-1.5B-Preview
Text Generation
• 2B • Updated • 1.64k
• • 74
agentica-org/DeepCoder-14B-Preview
Text Generation
• 15B • Updated • 585
• • 679
Feature Extraction
• 8B • Updated • 247
• 56
deepseek-ai/DeepSeek-R1-0528
Text Generation
• 685B • Updated • 456k
• • 2.46k
nvidia/Nemotron-Research-Reasoning-Qwen-1.5B
Text Generation
• 2B • Updated • 3.49k
• 244
Video-Text-to-Text
• 9B • Updated • 14
• 23
mistralai/Magistral-Small-2506
24B • Updated • 17.1k
• 610
microsoft/Phi-4-mini-reasoning
Text Generation
• 4B • Updated • 34.1k
• • 240
microsoft/Phi-4-mini-flash-reasoning
Text Generation
• 4B • Updated • 738
• 281
microsoft/Phi-4-reasoning
Text Generation
• 15B • Updated • 16.9k
• 228
osmosis-ai/Osmosis-Apply-1.7B
Text Generation
• 2B • Updated • 47
• 97
33B • Updated • 15
• 193
numind/NuMarkdown-8B-Thinking
Image-to-Text
• 8B • Updated • 14.4k
• 491
moonshotai/Kimi-K2-Thinking
Text Generation
• 1.1T • Updated • 85.6k
• • 1.71k
Text Generation
• 2B • Updated • 3.64k
• • 552
MaziyarPanahi/VibeThinker-1.5B-GGUF
Text Generation
• 2B • Updated • 1.73k
• 37
ServiceNow-AI/Apriel-1.5-15b-Thinker
Image-Text-to-Text
• 15B • Updated • 191
• 470
Image-Text-to-Text
• 10B • Updated • 415k
• • 619
Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled
Image-Text-to-Text
• 28B • Updated • 109k
• • 2.92k