Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 12 days ago • 106
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models Paper • 2606.11025 • Published Jun 9 • 41
Precision-RL Collection Defeating the Training-Inference Mismatch via FP16 • 2 items • Updated Nov 14, 2025
Precision-RL Collection Defeating the Training-Inference Mismatch via FP16 • 2 items • Updated Nov 14, 2025