LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures Paper • 2608.15242 • Published 16 days ago • 15
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Paper • 2606.16140 • Published Jun 15 • 126
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning Paper • 2603.21065 • Published Mar 22 • 79