You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

jev-9b

An unofficial JEV-style decision model trained with jev-torch. It reads a state, a question and a list of options, runs one forward pass, and returns a calibrated probability for every option.

Usage

pip install git+https://github.com/smha1012/jev-torch.git
from jev import JEVPredictor

jev = JEVPredictor("smha1012/jev-9b")
jev.predict(
    kind="noul",
    state="The canary shows p99 latency up 40% after the deploy.",
    question="Should the rollout be paused?",
    options=["false", "true"],
)

Question kinds: noul (exactly 2 options, false-like then true-like), choice (2–16 options), score (ordered levels 0–5).

Training

Evaluation

Agreement with the teacher

On the 25,376 rows of test_set_30k labelled by TypeSafe Jev 1.13, the same rows the autotrust/JEV cards report on. These measure how closely the model copies Jev, not accuracy on ground truth (see JevBench below).

Metric This model autotrust JEV-9B autotrust JEV-27B
Mean KL to teacher (lower is better) 0.0184 0.0190 0.0170
Choice top-1 agreement, teacher-labelled rows 90.1% 90.2% –
Choice top-1 agreement, all choice rows 89.9% 89.8% 90.3%
ECE vs teacher probabilities 0.0009 0.0007 0.0009

Per kind: choice 90.1%, noul 95.8%, score 87.9% top-1 agreement.

JevBench (Leanmcp v0.1)

Accuracy and calibration on 6,516 public benchmark cases with gold answers, next to TypeSafe Jev 1.13.0 on the same cases: 74.1% vs 84.9% accuracy (87.3% of Jev's), ECE 0.029 vs 0.020.

JevBench: jev-9b vs TypeSafe Jev 1.13.0

Limitations

  • Distilled, not RLCD. The model learns to reproduce TypeSafe Jev's output distributions (supervised KL + RPS) and is calibrated afterwards with temperature scaling. It does not reproduce Jev's own training method or architecture, which are unpublished, and cannot exceed the teacher.
  • Knowledge comes from the backbone. Gaps to Jev on JevBench are largest on knowledge-heavy questions (MMLU-Pro −24 points, MedQA −15) and did not shrink with training.
  • Long inputs. Agent trajectories longer than the 1,024-token training context trail Jev by ~14 points.
  • At most 16 options per choice question; trained mostly on short, synthetic English scenarios. Validate on your own domain before relying on it.

License and attribution

Weights: CC BY-NC 4.0, non-commercial use only. Most training labels are outputs of the closed TypeSafe Jev 1.13 model, so the weights are not offered for commercial use. For other uses, contact the author.

The code that trains and loads the model, jev-torch, is Apache-2.0. This is an independent project, not affiliated with TypeSafe AI or autotrust.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for smha1012/jev-9b

Finetuned
Qwen/Qwen3.5-9B
Adapter
(765)
this model

Dataset used to train smha1012/jev-9b