-
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
Paper • 2606.02060 • Published • 58 -
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
Paper • 2606.01993 • Published • 15 -
NJU-LINK/DR3-Eval
Viewer • Updated • 100 • 1.63k • 2 -
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation
Paper • 2606.02320 • Published • 15
Collections
Discover the best community collections!
Collections including paper arxiv:2606.01993
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100
-
Contrastive Decoding Improves Reasoning in Large Language Models
Paper • 2309.09117 • Published • 40 -
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Paper • 2310.08491 • Published • 57 -
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Paper • 2411.04282 • Published • 37 -
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Paper • 2411.14432 • Published • 26
-
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Paper • 2603.12056 • Published • 34 -
Memento-Skills: Let Agents Design Agents
Paper • 2603.18743 • Published • 58 -
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
Paper • 2603.15401 • Published • 20 -
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Paper • 2603.25158 • Published • 56
-
ControlLLM: Augment Language Models with Tools by Searching on Graphs
Paper • 2310.17796 • Published • 18 -
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster
Paper • 2311.08263 • Published • 16 -
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Paper • 2501.12599 • Published • 130 -
ARR: Question Answering with Large Language Models via Analyzing, Retrieving, and Reasoning
Paper • 2502.04689 • Published • 9
-
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
Paper • 2606.02060 • Published • 58 -
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
Paper • 2606.01993 • Published • 15 -
NJU-LINK/DR3-Eval
Viewer • Updated • 100 • 1.63k • 2 -
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation
Paper • 2606.02320 • Published • 15
-
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Paper • 2603.12056 • Published • 34 -
Memento-Skills: Let Agents Design Agents
Paper • 2603.18743 • Published • 58 -
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
Paper • 2603.15401 • Published • 20 -
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Paper • 2603.25158 • Published • 56
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100
-
ControlLLM: Augment Language Models with Tools by Searching on Graphs
Paper • 2310.17796 • Published • 18 -
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster
Paper • 2311.08263 • Published • 16 -
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Paper • 2501.12599 • Published • 130 -
ARR: Question Answering with Large Language Models via Analyzing, Retrieving, and Reasoning
Paper • 2502.04689 • Published • 9
-
Contrastive Decoding Improves Reasoning in Large Language Models
Paper • 2309.09117 • Published • 40 -
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Paper • 2310.08491 • Published • 57 -
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Paper • 2411.04282 • Published • 37 -
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Paper • 2411.14432 • Published • 26