Post
57
Vev: Jev-style decisions about images, from a 4B or 9B model you can run yourself.
Give it a screenshot or photo plus a yes/no, multiple-choice, or scoring question, and it returns a probability for every option instead of generating text.
It serves TypeSafe's
The clip shows
Try it in the browser:
CountingSheep/vev
Weights:
CountingSheep/vev-4b
CountingSheep/vev-9b
LoRA adapters are also available in the collection.
Code:
https://github.com/Xiaooolong/vev
Fine-tuned from Qwen3.5. Tested on NVIDIA GPUs so far.
Give it a screenshot or photo plus a yes/no, multiple-choice, or scoring question, and it returns a probability for every option instead of generating text.
It serves TypeSafe's
/v1/systemone format, so the official SDK works by changing the base URL, including requests with images.The clip shows
vev-4b playing Doom in real time. On each look, the image is split into 8 vertical slices, and Vev answers 8 yes/no questions in a single request, one per slice. A small fixed-rule harness turns those probabilities into turning and firing.Try it in the browser:
CountingSheep/vev
Weights:
CountingSheep/vev-4b
CountingSheep/vev-9b
LoRA adapters are also available in the collection.
Code:
https://github.com/Xiaooolong/vev
Fine-tuned from Qwen3.5. Tested on NVIDIA GPUs so far.