KABI
dongguanting
AI & ML interests
Reasoning and Alignment for Large Language Models
Recent Activity
liked a dataset about 5 hours ago
dongguanting/Agent-Reflex-RL-2K liked a model about 5 hours ago
dongguanting/Agent-Reflex-8B liked a dataset about 5 hours ago
dongguanting/Agent-Reflex-SFT-54KOrganizations
AEPO
The official datasets and model checkpoints of AEPO
-
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 109 -
dongguanting/Qwen3-8B-AEPO-DeepSearch
Text Generation • 8B • Updated • 23 • • 2 -
dongguanting/QwQ-32B-AEPO-DeepSearch
Text Generation • 33B • Updated • 24 • 2 -
dongguanting/Qwen3-14B-AEPO-DeepSearch
Robotics • 15B • Updated • 18 • 1
Tool-Star
Tool-Star is a reinforcement learning-based framework designed to empower LLMs to autonomously invoke multiple external tools during stepwise reasonin
-
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Paper • 2505.16410 • Published • 60 -
dongguanting/Tool-Star-SFT-54K
Viewer • Updated • 54k • 251 • 11 -
dongguanting/Multi-Tool-RL-10K
Viewer • Updated • 10k • 56 • 5 -
dongguanting/Tool-Star-Qwen-7B
Text Generation • 8B • Updated • 120 • 2
Agent-Reflex
AEPO
The official datasets and model checkpoints of AEPO
-
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 109 -
dongguanting/Qwen3-8B-AEPO-DeepSearch
Text Generation • 8B • Updated • 23 • • 2 -
dongguanting/QwQ-32B-AEPO-DeepSearch
Text Generation • 33B • Updated • 24 • 2 -
dongguanting/Qwen3-14B-AEPO-DeepSearch
Robotics • 15B • Updated • 18 • 1
ARPO
The official datasets and model checkpoints of ARPO
Tool-Star
Tool-Star is a reinforcement learning-based framework designed to empower LLMs to autonomously invoke multiple external tools during stepwise reasonin
-
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Paper • 2505.16410 • Published • 60 -
dongguanting/Tool-Star-SFT-54K
Viewer • Updated • 54k • 251 • 11 -
dongguanting/Multi-Tool-RL-10K
Viewer • Updated • 10k • 56 • 5 -
dongguanting/Tool-Star-Qwen-7B
Text Generation • 8B • Updated • 120 • 2
RAG-Critic