Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Abstract
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.
Community
This paper proposes GPD, a geometry-privileged distillation framework that routes question-relevant 3D cues and reference answers to a teacher during on-policy self-distillation, augmenting GRPO with distillation on incorrect trajectories while keeping inference RGB-only.
š» Code: https://github.com/ZJU-REAL/GPD
š¤ Model: https://huggingface.co/xinyili0624/GPD-4B
š¤ Dataset: https://huggingface.co/datasets/xinyili0624/GPD-15k
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation (2026)
- Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs (2026)
- SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning (2026)
- Distilling Visual Reasoning into Text Space (2026)
- GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning (2026)
- SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning (2026)
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 2
xinyili0624/GPD-2B
Datasets citing this paper 1
xinyili0624/GPD-15k
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper