LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation Paper • 2609.38146 • Published 10 days ago • 10
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis Paper • 2609.15309 • Published 25 days ago • 13
RECAP-Forcing: Retaining Content Appearances for Long Video Generation Paper • 2608.26671 • Published Aug 27 • 6
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning Paper • 2512.08135 • Published Dec 9, 2025
VideoNSA: Native Sparse Attention Scales Video Understanding Paper • 2510.02295 • Published Oct 2, 2025 • 10
OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps Paper • 2509.19282 • Published Sep 23, 2025 • 8
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation Paper • 2508.00728 • Published Aug 1, 2025
DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion Paper • 2507.22825 • Published Jul 30, 2025
Science-T2I: Addressing Scientific Illusions in Image Synthesis Paper • 2504.13129 • Published Apr 17, 2025 • 3
OmniControlNet: Dual-stage Integration for Conditional Image Generation Paper • 2406.05871 • Published Jun 9, 2024
RECAP-Forcing: Retaining Content Appearances for Long Video Generation Paper • 2608.26671 • Published Aug 27 • 6
PaintBench: Deterministic Evaluation of Precise Visual Editing Paper • 2606.00188 • Published May 29 • 4
Science-T2I: Addressing Scientific Illusions in Image Synthesis Paper • 2504.13129 • Published Apr 17, 2025 • 3
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs Paper • 2509.22646 • Published Sep 26, 2025 • 17
Transition Matching Distillation for Fast Video Generation Paper • 2601.09881 • Published Jan 14 • 34