OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use Paper • 2508.04482 • Published Aug 6, 2025 • 10
Exploring the Design Space of Reward Backpropagation for Flow Matching Paper • 2606.11075 • Published Jun 9 • 12
TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL Paper • 2606.28016 • Published Jun 26 • 1
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Paper • 2607.15273 • Published 14 days ago • 17
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published 9 days ago • 35
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Paper • 2607.15273 • Published 14 days ago • 17
TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL Paper • 2606.28016 • Published Jun 26 • 1
Exploring the Design Space of Reward Backpropagation for Flow Matching Paper • 2606.11075 • Published Jun 9 • 12
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching Paper • 2605.26108 • Published May 25 • 7
Fine-tuning Flow Matching Generative Models with Intermediate Feedback Paper • 2510.18072 • Published Oct 20, 2025
Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-training Paper • 2506.01376 • Published Jun 2, 2025
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use Paper • 2508.04482 • Published Aug 6, 2025 • 10
DecompOpt: Controllable and Decomposed Diffusion Models for Structure-based Molecular Optimization Paper • 2403.13829 • Published Mar 7, 2024
Global Sparse Momentum SGD for Pruning Very Deep Neural Networks Paper • 1909.12778 • Published Sep 27, 2019 • 1
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published 9 days ago • 35
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models Paper • 2606.11025 • Published Jun 9 • 41
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Paper • 2606.10968 • Published Jun 9 • 42
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Paper • 2606.10968 • Published Jun 9 • 42