Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs
Paper • 2610.05417 • Published • 1
Render2Reason (R2R) model from the paper Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs.
This model extends Qwen/Qwen3-VL-8B-Instruct with VGGT geometry features, fused into the visual tokens through gated cross-attention. It is trained jointly on spatial QA and an auxiliary novel-view semantic rendering task, using 32 input frames during training.
The model uses a custom architecture (Qwen3VLGeometryForConditionalGeneration), so it requires the code from the GitHub repository. Follow the installation steps there, then:
cd qwen-vl-finetune
hf download yuqun/R2R-Qwen3-VL-8B --local-dir checkpoints/R2R-Qwen3-VL-8B
python qwenvl/eval/demo.py \
--checkpoint checkpoints/R2R-Qwen3-VL-8B \
--video path/to/video.mp4 \
--question "How many chairs are in this room? Please answer the question using a single word or phrase."
The VGGT weights are loaded from facebook/VGGT-1B.
@article{wu2026render,
title = {Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs},
author = {Wu, Yuqun and Xiao, Yao and Zou, Chuhang and Wang, Shenlong and Hoiem, Derek},
journal = {arXiv preprint arXiv:2610.05417},
year = {2026}
}
Base model
Qwen/Qwen3-VL-8B-Instruct