Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks Paper • 2610.11794 • Published 4 days ago • 35
U-Space: Uncovering When and Why Uncertainty Arises in Language Models Paper • 2610.09087 • Published 6 days ago • 38
DecepEval: A Benchmark for Evaluating Deception in LLM Agents Paper • 2610.07967 • Published 6 days ago • 75
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement Paper • 2610.11959 • Published 4 days ago • 72
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 25 days ago • 228
Follow the Entities: A Corpus Map for Agentic Search Paper • 2609.37226 • Published 13 days ago • 100
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 16 days ago • 35
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 24 days ago • 138
NeoHorse-1 Collection NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness • 3 items • Updated Sep 9 • 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 321
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published Aug 31 • 63
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities Paper • 2608.28122 • Published Aug 28 • 66
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Paper • 2608.27454 • Published Aug 27 • 37
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published Aug 27 • 36
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 66