FaithEyes-8B-SFT
This is the cold-start supervised-fine-tuning (SFT) checkpoint of FaithEyes, trained from Qwen3-VL-8B-Instruct. It is the first stage of the two-stage SFT + RL pipeline in the FaithEyes framework.
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification
Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang
📖 Model Description
FaithEyes is a multi-agent self-judging framework for agentic vision-language models (VLMs). A single VLM plays two roles: a main agent that solves the visual question by interleaving reasoning with executable code-based tool calls, and a subagent that judges whether each process image (e.g. a crop) produced by the main agent is helpful for answering the question. The verdict {"is_helpful": true/false, "reasons": ...} is both injected into the tool observation to steer subsequent reasoning and used to scale the tool reward by the helpful-tool ratio, suppressing reward hacking.
This checkpoint is the SFT cold-start model. It equips Qwen3-VL-8B-Instruct with three capabilities prior to reinforcement learning:
- Code-based tool use — writing executable Python for image manipulation (cropping, zooming, rotation, contrast adjustment, arithmetic).
- Faithfulness judging — judging whether a process image helps answer the question and emitting the verdict as a JSON line.
- Feedback-driven reasoning — conditioning subsequent reasoning on the subagent's judgement rather than ignoring it.
The main agent and subagent share the same VLM and are trained jointly in this single stage, differing only in their system/user prompts. As reported in the paper, the SFT model merely imitates demonstrations without autonomously discriminating helpful from unhelpful tool calls; the RL checkpoint adds that discrimination ability.
📚 Citation
If you find this work useful, please consider citing:
@article{wang2026faitheyes,
title={FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification},
author={Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang},
journal={arXiv preprint arXiv:2607.28225},
year={2026}
}
🙏 Acknowledgements
This model is trained on top of Qwen3-VL-8B-Instruct. The training framework builds on verl, DeepEyes, Thyme, ms_swift, and VLMEvalKit.
- Downloads last month
- 25