NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale Paper • 2610.08430 • Published 5 days ago • 20
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms Paper • 2608.22335 • Published Aug 23 • 1
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published Sep 6 • 9
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published Sep 6 • 9
PA3: Policy-Aware Agent Alignment through Chain-of-Thought Paper • 2603.14602 • Published Mar 21 • 1
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms Paper • 2608.22335 • Published Aug 23 • 1
DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification Paper • 2605.27858 • Published May 27
Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention Paper • 2603.27187 • Published Mar 28