SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces Paper • 2610.11366 • Published 3 days ago • 12
TokenRouter: Efficient Serving System for Token-Level LLM Routing Paper • 2610.12242 • Published 3 days ago • 130
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement Paper • 2610.11959 • Published 3 days ago • 72
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 5 days ago • 91
SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting Paper • 2610.04109 • Published 9 days ago • 33
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 23 days ago • 80
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 23 days ago • 138
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 24 days ago • 228
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published Sep 10 • 232
OpenThinker-Agent2 Collection OpenThinker-Agent2: agentic SFT/RL datasets and 8B/32B models (cold-start SFT, RL, and the OpenThinkerAgent-32B release). • 11 items • Updated Jun 11 • 11
Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding Paper • 2606.21906 • Published Jun 20 • 27
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published Aug 31 • 63
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published Sep 7 • 36
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback Paper • 2608.13120 • Published Aug 13 • 32