nanovdr/ColNanoVDR-Q-Ettin150M-EVIE8B-4096-ML Sentence Similarity • 0.1B • Updated about 5 hours ago • 1
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 50
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents Paper • 2606.02031 • Published Jun 1 • 21
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published Jul 23 • 13
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published Jul 31 • 21
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding Paper • 2608.00036 • Published Jul 21 • 3
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Paper • 2607.04438 • Published Jul 5 • 62
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes Paper • 2607.04439 • Published Jul 5 • 64
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents Paper • 2606.31179 • Published Jun 30 • 7
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 66
CUAWright: A Minimal Unified Interface for Digital Agents Paper • 2610.04116 • Published 6 days ago • 5
Can Computation from Earlier Problems Help LLMs Solve New Ones? Paper • 2609.39394 • Published 8 days ago • 9
Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Paper • 2610.03509 • Published 6 days ago • 15
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers Paper • 2608.26762 • Published 12 days ago • 21