SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published 5 days ago • 127
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 13
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published 26 days ago • 79
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 116
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 96
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 123
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents Paper • 2602.07900 • Published Feb 8 • 4
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents Paper • 2602.07035 • Published Feb 3 • 31
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Paper • 2602.01785 • Published Feb 2 • 97
PaperBanana: Automating Academic Illustration for AI Scientists Paper • 2601.23265 • Published Jan 30 • 229
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents Paper • 2601.16746 • Published Jan 23 • 94
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning Paper • 2601.06943 • Published Jan 11 • 214
GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts Paper • 2601.05110 • Published Jan 8 • 29
Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents Paper • 2512.08870 • Published Dec 9, 2025 • 4
HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication Paper • 2510.10611 • Published Oct 12, 2025 • 5