-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 88 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 236 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 162 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
Collections
Discover the best community collections!
Collections including paper arxiv:2608.31046
-
Tuwhy/Qwen3-4B-OPSA
Text Generation • 4B • Updated • 992 • 1 -
Tuwhy/Qwen3.5-9B-OPSA
Image-Text-to-Text • 10B • Updated • 52 -
Tuwhy/Qwen3-1.7B-OPSA
Text Generation • 2B • Updated • 984 -
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
Paper • 2608.31046 • Published • 153
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 26 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 71 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 92 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 116
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 7 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 60 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 52
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 25 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 21 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 37 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 108
-
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 521 -
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 555 -
LightMem: Lightweight and Efficient Memory-Augmented Generation
Paper • 2510.18866 • Published • 117 -
The End of Manual Decoding: Towards Truly End-to-End Language Models
Paper • 2510.26697 • Published • 121
-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 88 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 236 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 162 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 7 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 60 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 52
-
Tuwhy/Qwen3-4B-OPSA
Text Generation • 4B • Updated • 992 • 1 -
Tuwhy/Qwen3.5-9B-OPSA
Image-Text-to-Text • 10B • Updated • 52 -
Tuwhy/Qwen3-1.7B-OPSA
Text Generation • 2B • Updated • 984 -
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
Paper • 2608.31046 • Published • 153
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 25 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 21 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 37 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 108
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 26 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 71 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 92 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 116
-
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 521 -
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 555 -
LightMem: Lightweight and Efficient Memory-Augmented Generation
Paper • 2510.18866 • Published • 117 -
The End of Manual Decoding: Towards Truly End-to-End Language Models
Paper • 2510.26697 • Published • 121