SanSi: A Looped Typed Decision Model for System 1.5 Thinking Paper • 2610.07730 • Published 5 days ago • 13
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym Paper • 2609.37267 • Published 12 days ago • 40
When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models Paper • 2610.05719 • Published 6 days ago • 32
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 11 days ago • 61
Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model Paper • 2609.40358 • Published 11 days ago • 21
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 11 days ago • 64
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 12 days ago • 64
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 11 days ago • 120
Game-Guided Skill Discovery through Self-Play for Playable Agent Control Paper • 2609.40137 • Published 11 days ago • 6
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 11 days ago • 111
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 12 days ago • 43
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 12 days ago • 65
Follow the Entities: A Corpus Map for Agentic Search Paper • 2609.37226 • Published 12 days ago • 100
Recursive Harness Distillation across Agents for Robot Manipulation Paper • 2609.33378 • Published 14 days ago • 42
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 14 days ago • 46
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge Paper • 2609.34327 • Published 13 days ago • 41
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 25 days ago • 58
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 25 days ago • 67