CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-prm-eta100-Advorm-stepsplit-none 2B • Updated 4 days ago • 51
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-prm-eta100-Advorm-stepsplit-none 2B • Updated 4 days ago • 51
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-grpo-useKL_True-KLlossCoef1e-3 2B • Updated 4 days ago • 163
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-grpo-useKL_True-KLlossCoef1e-3 2B • Updated 4 days ago • 163
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published 8 days ago • 86
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents Paper • 2412.13549 • Published Dec 18, 2024
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving Paper • 2510.11769 • Published Oct 13, 2025 • 26
ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning Paper • 2510.12693 • Published Oct 14, 2025 • 28
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Paper • 2603.13985 • Published Mar 14 • 10
AgentSPEX: An Agent SPecification and EXecution Language Paper • 2604.13346 • Published 24 days ago • 162
AgentSPEX: An Agent SPecification and EXecution Language Paper • 2604.13346 • Published 24 days ago • 162
Seedance 2.0: Advancing Video Generation for World Complexity Paper • 2604.14148 • Published 23 days ago • 155