FaithEyes-8B-SFT

This is the cold-start supervised-fine-tuning (SFT) checkpoint of FaithEyes, trained from Qwen3-VL-8B-Instruct. It is the first stage of the two-stage SFT + RL pipeline in the FaithEyes framework.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification

Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang

arXiv Hugging Face ModelScope GitHub


📖 Model Description

FaithEyes is a multi-agent self-judging framework for agentic vision-language models (VLMs). A single VLM plays two roles: a main agent that solves the visual question by interleaving reasoning with executable code-based tool calls, and a subagent that judges whether each process image (e.g. a crop) produced by the main agent is helpful for answering the question. The verdict {"is_helpful": true/false, "reasons": ...} is both injected into the tool observation to steer subsequent reasoning and used to scale the tool reward by the helpful-tool ratio, suppressing reward hacking.

This checkpoint is the SFT cold-start model. It equips Qwen3-VL-8B-Instruct with three capabilities prior to reinforcement learning:

  1. Code-based tool use — writing executable Python for image manipulation (cropping, zooming, rotation, contrast adjustment, arithmetic).
  2. Faithfulness judging — judging whether a process image helps answer the question and emitting the verdict as a JSON line.
  3. Feedback-driven reasoning — conditioning subsequent reasoning on the subagent's judgement rather than ignoring it.

The main agent and subagent share the same VLM and are trained jointly in this single stage, differing only in their system/user prompts. As reported in the paper, the SFT model merely imitates demonstrations without autonomously discriminating helpful from unhelpful tool calls; the RL checkpoint adds that discrimination ability.

📚 Citation

If you find this work useful, please consider citing:

@article{wang2026faitheyes,
  title={FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification},
  author={Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang},
  journal={arXiv preprint arXiv:2607.28225},
  year={2026}
}

🙏 Acknowledgements

This model is trained on top of Qwen3-VL-8B-Instruct. The training framework builds on verl, DeepEyes, Thyme, ms_swift, and VLMEvalKit.

Downloads last month
25
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jackwang111/FaithEyes-8B-SFT

Finetuned
(616)
this model
Quantizations
1 model

Collection including Jackwang111/FaithEyes-8B-SFT

Paper for Jackwang111/FaithEyes-8B-SFT