Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 6 days ago • 524
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published Aug 7 • 107
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 311
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published Jul 19 • 167
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 62
Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs Paper • 2505.11277 • Published May 16, 2025 • 64
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published Jul 14 • 20
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 86
AlayaWorld: Long-Horizon and Playable Video World Generation Paper • 2607.06291 • Published Jul 7 • 92
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark Paper • 2606.30170 • Published Jun 29 • 8
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Paper • 2606.24428 • Published Jun 23 • 53
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 178
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams Paper • 2606.21337 • Published Jun 19 • 76
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 254
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 142
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 173
Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Paper • 2606.08063 • Published Jun 6 • 82
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 108