R2R-Qwen3-VL-8B

Render2Reason (R2R) model from the paper Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs.

[Project Page] [Paper] [Code]

This model extends Qwen/Qwen3-VL-8B-Instruct with VGGT geometry features, fused into the visual tokens through gated cross-attention. It is trained jointly on spatial QA and an auxiliary novel-view semantic rendering task, using 32 input frames during training.

Usage

The model uses a custom architecture (Qwen3VLGeometryForConditionalGeneration), so it requires the code from the GitHub repository. Follow the installation steps there, then:

cd qwen-vl-finetune
hf download yuqun/R2R-Qwen3-VL-8B --local-dir checkpoints/R2R-Qwen3-VL-8B
python qwenvl/eval/demo.py \
    --checkpoint checkpoints/R2R-Qwen3-VL-8B \
    --video path/to/video.mp4 \
    --question "How many chairs are in this room? Please answer the question using a single word or phrase."

The VGGT weights are loaded from facebook/VGGT-1B.

Citation

@article{wu2026render,
  title   = {Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs},
  author  = {Wu, Yuqun and Xiao, Yao and Zou, Chuhang and Wang, Shenlong and Hoiem, Derek},
  journal = {arXiv preprint arXiv:2610.05417},
  year    = {2026}
}
Downloads last month
35
Safetensors
Model size
10B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yuqun/R2R-Qwen3-VL-8B

Finetuned
(634)
this model

Space using yuqun/R2R-Qwen3-VL-8B 1

Paper for yuqun/R2R-Qwen3-VL-8B