Instructions to use smha1012/jev-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use smha1012/jev-9b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
jev-9b
An unofficial JEV-style decision model trained with jev-torch. It reads a state, a question and a list of options, runs one forward pass, and returns a calibrated probability for every option.
Usage
pip install git+https://github.com/smha1012/jev-torch.git
from jev import JEVPredictor
jev = JEVPredictor("smha1012/jev-9b")
jev.predict(
kind="noul",
state="The canary shows p99 latency up 40% after the deploy.",
question="Should the rollout be paused?",
options=["false", "true"],
)
Question kinds: noul (exactly 2 options, false-like then true-like), choice (2–16 options),
score (ordered levels 0–5).
Training
- Training data: jev_distill (https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3 @ fc99c6357a9f89f7512c4a987314352addead049)
- Base model:
Qwen/Qwen3.5-9B(frozen) + LoRA r=16, α=32 + 24-slot fp32 decision head - Loss:
1·kl + 0.5·rps[score] - Optimization: 4750 steps × 128 rows, LoRA lr 0.0001, head lr 0.0002
- Temperatures: noul 1.008, choice 1.000, score 1.008
Evaluation
Agreement with the teacher
On the 25,376 rows of test_set_30k labelled by TypeSafe Jev 1.13, the same rows the autotrust/JEV cards
report on. These measure how closely the model copies Jev, not accuracy on ground truth (see JevBench below).
| Metric | This model | autotrust JEV-9B | autotrust JEV-27B |
|---|---|---|---|
| Mean KL to teacher (lower is better) | 0.0184 | 0.0190 | 0.0170 |
| Choice top-1 agreement, teacher-labelled rows | 90.1% | 90.2% | – |
| Choice top-1 agreement, all choice rows | 89.9% | 89.8% | 90.3% |
| ECE vs teacher probabilities | 0.0009 | 0.0007 | 0.0009 |
Per kind: choice 90.1%, noul 95.8%, score 87.9% top-1 agreement.
JevBench (Leanmcp v0.1)
Accuracy and calibration on 6,516 public benchmark cases with gold answers, next to TypeSafe Jev 1.13.0 on the same cases: 74.1% vs 84.9% accuracy (87.3% of Jev's), ECE 0.029 vs 0.020.
Limitations
- Distilled, not RLCD. The model learns to reproduce TypeSafe Jev's output distributions (supervised KL + RPS) and is calibrated afterwards with temperature scaling. It does not reproduce Jev's own training method or architecture, which are unpublished, and cannot exceed the teacher.
- Knowledge comes from the backbone. Gaps to Jev on JevBench are largest on knowledge-heavy questions (MMLU-Pro −24 points, MedQA −15) and did not shrink with training.
- Long inputs. Agent trajectories longer than the 1,024-token training context trail Jev by ~14 points.
- At most 16 options per choice question; trained mostly on short, synthetic English scenarios. Validate on your own domain before relying on it.
License and attribution
Weights: CC BY-NC 4.0, non-commercial use only. Most training labels are outputs of the closed TypeSafe Jev 1.13 model, so the weights are not offered for commercial use. For other uses, contact the author.
The code that trains and loads the model, jev-torch, is Apache-2.0. This is an independent project, not affiliated with TypeSafe AI or autotrust.
- Downloads last month
- -
