How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published 19 days ago • 48
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning Paper • 2607.08940 • Published Jul 9 • 2
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Paper • 2606.28186 • Published Jun 26 • 9
Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models Paper • 2606.16700 • Published Jun 15 • 15
Guava: An Effective and Universal Harness for Embodied Manipulation Paper • 2606.18363 • Published Jun 16 • 28
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers Paper • 2606.00579 • Published May 30 • 2
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Paper • 2606.06574 • Published Jun 4 • 25
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents Paper • 2604.18543 • Published Apr 20 • 30
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook Paper • 2602.14299 • Published Feb 15 • 27
What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis Paper • 2602.12395 • Published Feb 12 • 17
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models Paper • 2602.02140 • Published Feb 2 • 12
TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models Paper • 2601.18744 • Published Jan 26 • 11
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models Paper • 2512.19995 • Published Dec 23, 2025 • 16
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction Paper • 2512.18880 • Published Dec 21, 2025 • 25
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions Paper • 2512.11995 • Published Dec 12, 2025 • 10
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs Paper • 2511.07419 • Published Nov 10, 2025 • 27