JevK5-9B v0.3.3 — open-weight Jev alternative, 9B

JevK5-9B is the larger sibling of JevK5: the same v0.3 training data on Qwen3.5-9B. It is an independent, Apache-2.0 open-source alternative to TypeSafe's Jev for typed decisions. It reads a state and a yes/no (noul), choice, or score question and returns a probability for every option in one forward pass, with zero generated tokens. It runs on the JevK5 runtime, which also serves a TypeSafe-style /v1/systemone endpoint. This is not Jev's model or architecture and is not affiliated with TypeSafe AI.

v0.3.3 replaces v0.3's 9B weights. The data is unchanged; training ran for 1.5 epochs instead of 1, chosen by a sweep with a rule fixed before the results (below). v0.3's 9B stays available under the Hub tag v0.3.

Use the JevK5 runtime shown below to read option probabilities. Generic text-generation examples on the Hub call generate() and do not perform JevK5's decision readout.

  • Base: Qwen3.5-9B, with a LoRA (rank 16, attention projections) merged into the weights
  • Readout: SemIf's protocol (TheoLeeCJ/SemIf, MIT): a softmax over the answer letters' next-token logits, divided by one calibration temperature. Questions with more than 16 options are read in several passes and combined with a second temperature. Both are in jevk5_config.json: temperature 1.316 and knockout_temperature 1.05
  • Runtime: github.com/allebee/jevk5, 0.3.0 or later, with CUDA graphs: about 16 ms per short decision on an otherwise idle H100. Needs a CUDA GPU with about 19 GB for bf16
  • License: Apache-2.0

When to use the 9B

Choose the 4B when speed or memory matters: it needs about 9 GB and is about 20% faster. The 9B v0.3.3 is ahead of it on our held-out checks and ties it on JevBench's public items.

  • Held-out checks, 9B v0.3.3 against 4B: index proxy 0.794 against 0.731, teacher questions 0.848 against 0.834, hand-written hard set 0.859 against 0.766 (55 against 49 of 64). It also leads by 8-10 points on questions with more than 16 options, and on bev-decision's test sample (0.698 against 0.663).
  • JevBench's public items, 9B v0.3.3 against 4B: hard tier 0.775 against 0.784 (7 items fixed, 8 broken), standard tier 0.958 against 0.944. Hard-tier calibration is still worse than the 4B's (ECE 0.071 against 0.054).

Other sizes and formats

  • JevK5 (4B, v0.3): the same data on Qwen3.5-4B, ~9 GB in bf16.
  • JevK5-GGUF: GGUF builds for llama.cpp (NVIDIA, AMD, Intel and Apple GPUs, or a CPU). The 9B v0.3.3 Q8_0 gives the same answer as these weights on 229 of 231 public JevBench items, and the Q5_K_M on 228. A 9B v0.3.3 Q4_K_M agreed on 222 but lost 5 hard items (0.775 → 0.730) and is not published.

Why 1.5 epochs: a pre-registered sweep

v0.3 trained both sizes for one epoch. A size control (the same 1-epoch data with 20k rows repeated) gained index proxy without new content, which suggested undertraining, so we trained full runs (each with its own learning-rate schedule, same data, seed and dev set) and fixed the rule first: per size, take the epoch count with the highest mean of index proxy, teacher accuracy and hand-written hard set; adopt it only if that mean beats 1 epoch by at least 0.02 and none of the three drops by more than 0.03.

Model Epochs Index proxy Teacher questions Hard set Mean Verdict
9B 1 (v0.3) 0.762 0.851 0.781 0.798
9B 1.5 0.794 0.848 0.859 0.834 adopted: this release
9B 2 0.784 0.865 0.828 0.826
4B 1 (v0.3) 0.732 0.834 0.766 0.777 kept
4B 2 0.769 0.826 0.766 0.787 below the +0.02 bar

This is one seed per setting, and each index source in the dev set has about 40 rows.

Results

Held-out checks (used to choose between models)

The dev set is held-out rows only: teacher questions from three domains that training never saw (residential leases, public-sector permits, manufacturing QC), a hand-written hard set, and a hashed 5% of every public train split. All columns are read the same way, through the runtime.

JevK5 v0.2 (4B) JevK5 v0.3 (4B) JevK5-9B v0.3 JevK5-9B v0.3.3
Index proxy (16 sources, see below) 0.620 0.731 0.762 0.794
Held-out teacher questions (362), accuracy 0.801 0.834 0.851 0.848
Hand-written hard set (64), accuracy 0.766 0.766 0.781 0.859
ECE on the teacher questions, calibrated 0.034 0.035 0.032 0.043

The index proxy is our own estimate, not an index score: the chance-corrected skill averaged over the 16 dev sources that are held-out train-split rows of Jev Decision Index benchmarks, 40 rows each, so each source alone is noisy (about ±0.15). Against v0.3's 9B, v0.3.3 gains most on NLI4CT (0.55 → 0.70), WinoGrande (0.85 → 0.95) and Amazon ESCI (0.53 → 0.60); its largest loss is HellaSwag (0.87 → 0.83). Calibration on the teacher questions is a little worse (ECE 0.043 against 0.032).

bev-decision-150K (another group's decision mix, evaluation only)

On 2,500 hashed rows of the test split of avbiswas/bev-decision-150K (4,723 questions, every question type): v0.3.3 scores 0.698 (v0.3's 9B 0.700, the v0.3 4B 0.663), with ECE 0.033 (v0.3's 9B 0.038). By type: choice 0.732, yes/no 0.774, score 0.482. Leaving out the 34 questions whose document also appears in our training data changes accuracy by 0.002.

JevBench v1.2 public items (report only)

231 public items through JevBench's own runner (jevk5_direct adapter): 231/231 valid, 0 failures.

Split n JevK5 v0.3 (4B) JevK5-9B v0.3 JevK5-9B v0.3.3 v0.3.3 ECE (4B)
easy 48 1.000 1.000 1.000 0.016 (0.018)
original (standard) 72 0.944 0.944 0.958 0.033 (0.057)
hard (public half) 111 0.784 0.730 0.775 0.071 (0.054)
  • Against v0.3's 9B: 7 items fixed and 1 broken (6 and 1 on the hard tier; McNemar p = 0.07 overall). Hard-tier ECE 0.071 against 0.126; distance to the exact gold distributions on the 10 probability items 0.251 against 0.284.
  • Against the v0.3 4B: 9 fixed and 9 broken (p = 1.0); hard-tier ECE 0.071 against 0.054.
  • By family on the hard tier, v0.3 9B → v0.3.3: probability 0.70 → 0.90, long policies 0.58 → 0.74, ambiguous 0.71 → 0.86, multi-hop 0.83 → 0.78. Dates and numbers stay weak (0.47).
  • Latency (H100, in-process, batch 1, CUDA graphs, GPU otherwise idle): p50 16 ms, p95 18 ms on easy and standard items; hard items p50 42 ms, p95 223 ms. v0.3's 9B card reported 31 ms measured while another job shared the GPU; the model size is the same.

More than 16 options

These use the runtime's knockout readout (groups of up to 16, then a final) at this repo's knockout_temperature. The runs are 500 train-split items per dataset in the Decision Index's request shape, with every option offered. None of these items is in the training or dev data.

Train split Options JevK5 v0.3 (4B) accuracy / ECE JevK5-9B v0.3 accuracy / ECE JevK5-9B v0.3.3 accuracy / ECE v0.3.3 macro-F1
MASSIVE en-US (fitting set) 60 0.738 / 0.045 0.818 / 0.036 0.814 / 0.031 0.809
BANKING77 77 0.652 / 0.044 0.734 / 0.044 0.754 / 0.055 0.739
CLINC150 with out-of-scope 151 0.700 / 0.056 0.780 / 0.061 0.804 / 0.060 0.819
  • On CLINC150, out-of-scope recall is 0.59 (v0.3's 9B 0.46) with precision 0.84 (0.76).

  • Second temperature: v0.3.3 carries its own, 1.05, fitted by NLL on the MASSIVE items only (the two halves give 1.07 and 1.02). With the old 0.77 it would be overconfident (MASSIVE ECE 0.073, BANKING77 0.114).

Jev Decision Index and JevBench

JevK5-9B has not been run by the Jev Decision Index or submitted to JevBench yet. For reference, the index reran JevK5 0.2.2 (v0.2 4B weights, with the runtime that answers any number of options) on its previously refused rows (discussion #13). It scored 36.31, 15th of 49, and its calibration was 4th best (ECE 0.031). JevBench v1.4 ranked JevK5 v0.2 #2 of 76 systems.

How it was trained

On the same 47,460 rows as JevK5 v0.3 (4B); see that card for the full data table with licenses.

  • Teacher questions (17,408): 3,270 from Qwen3.6-27B (Apache-2.0, self-hosted, thinking on) and 14,138 from GPT-6 Luna (OpenAI, through OpenAI's API). Each question was answered twice, independently, by its teacher and kept only when both answers matched the intended one. Luna's outputs were generated under OpenAI's terms, which govern their use; review them for your use case.
  • Public replay (30,052 items): train splits of 26 public datasets, with no test or validation split of any dataset, and no split of MMLU, MMLU-Pro, ANLI or NLI4CT. Each dataset's license is in the JevK5 card's table.
  • Training: cross-entropy on the option-letter logits, SemIf's prompt format, 1.5 epochs, learning rate 3e-5, inputs up to 2,048 tokens. One temperature (1.316) is fitted on the held-out teacher questions. The second temperature (1.05) is fitted by NLL on the 500 MASSIVE items above.
  • Not released: the 2-epoch 9B (higher on teacher questions, 0.865, but lower on the index proxy and the hard set; the pre-registered rule picked 1.5 epochs), and the two development 9Bs that trained on ANLI/NLI4CT or on MMLU's auxiliary_train (RACE, non-commercial), as described on the v0.3 card.

Data rules.

  • No JevBench item, public or held out, and no output of Jev was used for training, tuning or selection. JevBench's public items were only used to report the numbers above.
  • Every public and Luna row was checked against the test and validation text of 35 Decision Index benchmarks, and against JevBench's public items. A row was dropped for an exact match or for any shared 8-word sequence. GPQA and HLE are gated and were not checked; no source is built from them.

Declared overlap with the Jev Decision Index. These are train splits of index benchmarks, deduplicated against their test and validation items: ARC, OpenBookQA, CommonsenseQA, GSM8K, WinoGrande, HellaSwag, BANKING77, CLINC150, SGD, Amazon ESCI, When2Call, iSarcasmEval, RAGTruth, HoVer and the New Yorker caption contest. Separately, 53 of the Qwen-written training questions share at least one 8-word sequence with ContractNLI (41) or SGD (12) test or dev text; they were reported by the scan and kept.

Known weak spots

  • One seed, small per-source dev slices. The sweep's winner rests on one run per setting, and each index source has about 40 dev rows.
  • Calibration on the held-out teacher questions is a little worse than v0.3's 9B (ECE 0.043 against 0.032), and hard-tier calibration on JevBench is worse than the 4B's (0.071 against 0.054).
  • On the dev set's teacher families, trap (0.76 → 0.64) and ambiguous (0.75 → 0.67) questions dropped against v0.3's 9B while judging (0.83 → 0.91) and dates and numbers (0.32 → 0.53) rose.
  • Dates and numbers stay weak on JevBench's hard tier (0.47). On bev-decision, score questions are the weakest type (0.482), and BANKING77's calibration over 77 options is a little worse than v0.3's 9B (ECE 0.055 against 0.044).
  • Accuracy drops sharply on JevBench's fresh sealed decisions (measured for JevK5 v0.2). Real-world workflow performance against Jev has not been measured.
  • English only. Needs a CUDA GPU with ~19 GB for bf16. Inputs over 16,384 tokens are refused, not cut.

Use

from jevk5 import JevK5

model = JevK5("alibiserikbay/JevK5-9B")
model.decide(
    "I was billed twice for order #4411. Please refund the duplicate charge today.",
    {"type": "choice", "instructions": "Which team should handle this?",
     "criteria": {"billing": "Payments and refunds", "tech": "Bugs", "sales": "New purchases"}},
)
# {'type': 'choice', 'confidence': 0.993, 'choice': 'billing', 'probabilities': {...}, ...}

With llama.cpp (see JevK5-GGUF):

from jevk5 import JevK5GGUF

model = JevK5GGUF("http://127.0.0.1:8080", temperature=1.316, knockout_temperature=1.05)

Or as a server that answers TypeSafe-style /v1/systemone requests: jevk5-serve --model alibiserikbay/JevK5-9B --port 8090. Both read temperature and knockout_temperature from this repo's jevk5_config.json.

Credits

Qwen3.5-9B and Qwen3.6-27B by the Qwen team (Apache-2.0). GPT-6 Luna by OpenAI. The one-pass readout and prompt come from SemIf by TheoLeeCJ (MIT). The public datasets listed above belong to their authors, under their licenses (table on the JevK5 card). Evaluated with JevBench (github.com/fstandhartinger/jevbench, MIT) and bev-decision-150K. Not affiliated with TypeSafe AI or Jev.

Downloads last month
1,022
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alibiserikbay/JevK5-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(991)
this model
Finetunes
1 model