SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 10 days ago • 266
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 17 days ago • 183
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Paper • 2608.09873 • Published Aug 10 • 29
Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Paper • 2605.04128 • Published May 5 • 18
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 139
NeuroCogMap Reveals Cognitive Organization of Large Language Models Paper • 2607.00397 • Published Jul 1 • 8
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 164
DiffusionBench: On Holistic Evaluation of Diffusion Transformers Paper • 2606.24888 • Published Jun 23 • 12
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning Paper • 2606.24548 • Published Jun 23 • 12
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective Paper • 2605.28819 • Published May 27 • 8
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Paper • 2605.21573 • Published May 20 • 109
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Paper • 2605.21468 • Published May 20 • 51
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199