Instructions to use alibiserikbay/JevK5-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alibiserikbay/JevK5-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="alibiserikbay/JevK5-9B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("alibiserikbay/JevK5-9B") model = AutoModelForCausalLM.from_pretrained("alibiserikbay/JevK5-9B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use alibiserikbay/JevK5-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alibiserikbay/JevK5-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alibiserikbay/JevK5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/alibiserikbay/JevK5-9B
- SGLang
How to use alibiserikbay/JevK5-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "alibiserikbay/JevK5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alibiserikbay/JevK5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "alibiserikbay/JevK5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alibiserikbay/JevK5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use alibiserikbay/JevK5-9B with Docker Model Runner:
docker model run hf.co/alibiserikbay/JevK5-9B
JevK5-9B v0.3.3 — open-weight Jev alternative, 9B
JevK5-9B is the larger sibling of JevK5: the same
v0.3 training data on Qwen3.5-9B. It is an independent, Apache-2.0 open-source alternative to
TypeSafe's Jev for typed decisions. It reads a state and a yes/no (noul), choice, or score question
and returns a probability for every option in one forward pass, with zero generated tokens. It
runs on the JevK5 runtime, which also serves a TypeSafe-style
/v1/systemone endpoint. This is not Jev's model or architecture and is not affiliated with TypeSafe
AI.
v0.3.3 replaces v0.3's 9B weights. The data is unchanged; training ran for 1.5 epochs instead
of 1, chosen by a sweep with a rule fixed before the results (below). v0.3's 9B stays available
under the Hub tag v0.3.
Use the JevK5 runtime shown below to read option probabilities. Generic text-generation examples
on the Hub call generate() and do not perform JevK5's decision readout.
- Base: Qwen3.5-9B, with a LoRA (rank 16, attention projections) merged into the weights
- Readout: SemIf's protocol (TheoLeeCJ/SemIf, MIT): a softmax over the answer letters'
next-token logits, divided by one calibration temperature. Questions with more than 16 options
are read in several passes and combined with a second temperature. Both are in
jevk5_config.json:temperature1.316 andknockout_temperature1.05 - Runtime: github.com/allebee/jevk5, 0.3.0 or later, with CUDA graphs: about 16 ms per short decision on an otherwise idle H100. Needs a CUDA GPU with about 19 GB for bf16
- License: Apache-2.0
When to use the 9B
Choose the 4B when speed or memory matters: it needs about 9 GB and is about 20% faster. The 9B v0.3.3 is ahead of it on our held-out checks and ties it on JevBench's public items.
- Held-out checks, 9B v0.3.3 against 4B: index proxy 0.794 against 0.731, teacher questions 0.848 against 0.834, hand-written hard set 0.859 against 0.766 (55 against 49 of 64). It also leads by 8-10 points on questions with more than 16 options, and on bev-decision's test sample (0.698 against 0.663).
- JevBench's public items, 9B v0.3.3 against 4B: hard tier 0.775 against 0.784 (7 items fixed, 8 broken), standard tier 0.958 against 0.944. Hard-tier calibration is still worse than the 4B's (ECE 0.071 against 0.054).
Other sizes and formats
- JevK5 (4B, v0.3): the same data on Qwen3.5-4B, ~9 GB in bf16.
- JevK5-GGUF: GGUF builds for llama.cpp (NVIDIA, AMD, Intel and Apple GPUs, or a CPU). The 9B v0.3.3 Q8_0 gives the same answer as these weights on 229 of 231 public JevBench items, and the Q5_K_M on 228. A 9B v0.3.3 Q4_K_M agreed on 222 but lost 5 hard items (0.775 → 0.730) and is not published.
Why 1.5 epochs: a pre-registered sweep
v0.3 trained both sizes for one epoch. A size control (the same 1-epoch data with 20k rows repeated) gained index proxy without new content, which suggested undertraining, so we trained full runs (each with its own learning-rate schedule, same data, seed and dev set) and fixed the rule first: per size, take the epoch count with the highest mean of index proxy, teacher accuracy and hand-written hard set; adopt it only if that mean beats 1 epoch by at least 0.02 and none of the three drops by more than 0.03.
| Model | Epochs | Index proxy | Teacher questions | Hard set | Mean | Verdict |
|---|---|---|---|---|---|---|
| 9B | 1 (v0.3) | 0.762 | 0.851 | 0.781 | 0.798 | |
| 9B | 1.5 | 0.794 | 0.848 | 0.859 | 0.834 | adopted: this release |
| 9B | 2 | 0.784 | 0.865 | 0.828 | 0.826 | |
| 4B | 1 (v0.3) | 0.732 | 0.834 | 0.766 | 0.777 | kept |
| 4B | 2 | 0.769 | 0.826 | 0.766 | 0.787 | below the +0.02 bar |
This is one seed per setting, and each index source in the dev set has about 40 rows.
Results
Held-out checks (used to choose between models)
The dev set is held-out rows only: teacher questions from three domains that training never saw (residential leases, public-sector permits, manufacturing QC), a hand-written hard set, and a hashed 5% of every public train split. All columns are read the same way, through the runtime.
| JevK5 v0.2 (4B) | JevK5 v0.3 (4B) | JevK5-9B v0.3 | JevK5-9B v0.3.3 | |
|---|---|---|---|---|
| Index proxy (16 sources, see below) | 0.620 | 0.731 | 0.762 | 0.794 |
| Held-out teacher questions (362), accuracy | 0.801 | 0.834 | 0.851 | 0.848 |
| Hand-written hard set (64), accuracy | 0.766 | 0.766 | 0.781 | 0.859 |
| ECE on the teacher questions, calibrated | 0.034 | 0.035 | 0.032 | 0.043 |
The index proxy is our own estimate, not an index score: the chance-corrected skill averaged over the 16 dev sources that are held-out train-split rows of Jev Decision Index benchmarks, 40 rows each, so each source alone is noisy (about ±0.15). Against v0.3's 9B, v0.3.3 gains most on NLI4CT (0.55 → 0.70), WinoGrande (0.85 → 0.95) and Amazon ESCI (0.53 → 0.60); its largest loss is HellaSwag (0.87 → 0.83). Calibration on the teacher questions is a little worse (ECE 0.043 against 0.032).
bev-decision-150K (another group's decision mix, evaluation only)
On 2,500 hashed rows of the test split of avbiswas/bev-decision-150K (4,723 questions, every question type): v0.3.3 scores 0.698 (v0.3's 9B 0.700, the v0.3 4B 0.663), with ECE 0.033 (v0.3's 9B 0.038). By type: choice 0.732, yes/no 0.774, score 0.482. Leaving out the 34 questions whose document also appears in our training data changes accuracy by 0.002.
JevBench v1.2 public items (report only)
231 public items through JevBench's own runner (jevk5_direct adapter): 231/231 valid, 0 failures.
| Split | n | JevK5 v0.3 (4B) | JevK5-9B v0.3 | JevK5-9B v0.3.3 | v0.3.3 ECE (4B) |
|---|---|---|---|---|---|
| easy | 48 | 1.000 | 1.000 | 1.000 | 0.016 (0.018) |
| original (standard) | 72 | 0.944 | 0.944 | 0.958 | 0.033 (0.057) |
| hard (public half) | 111 | 0.784 | 0.730 | 0.775 | 0.071 (0.054) |
- Against v0.3's 9B: 7 items fixed and 1 broken (6 and 1 on the hard tier; McNemar p = 0.07 overall). Hard-tier ECE 0.071 against 0.126; distance to the exact gold distributions on the 10 probability items 0.251 against 0.284.
- Against the v0.3 4B: 9 fixed and 9 broken (p = 1.0); hard-tier ECE 0.071 against 0.054.
- By family on the hard tier, v0.3 9B → v0.3.3: probability 0.70 → 0.90, long policies 0.58 → 0.74, ambiguous 0.71 → 0.86, multi-hop 0.83 → 0.78. Dates and numbers stay weak (0.47).
- Latency (H100, in-process, batch 1, CUDA graphs, GPU otherwise idle): p50 16 ms, p95 18 ms on easy and standard items; hard items p50 42 ms, p95 223 ms. v0.3's 9B card reported 31 ms measured while another job shared the GPU; the model size is the same.
More than 16 options
These use the runtime's knockout readout (groups of up to 16, then a final) at this repo's
knockout_temperature. The runs are 500 train-split items per dataset in the Decision Index's request
shape, with every option offered. None of these items is in the training or dev data.
| Train split | Options | JevK5 v0.3 (4B) accuracy / ECE | JevK5-9B v0.3 accuracy / ECE | JevK5-9B v0.3.3 accuracy / ECE | v0.3.3 macro-F1 |
|---|---|---|---|---|---|
| MASSIVE en-US (fitting set) | 60 | 0.738 / 0.045 | 0.818 / 0.036 | 0.814 / 0.031 | 0.809 |
| BANKING77 | 77 | 0.652 / 0.044 | 0.734 / 0.044 | 0.754 / 0.055 | 0.739 |
| CLINC150 with out-of-scope | 151 | 0.700 / 0.056 | 0.780 / 0.061 | 0.804 / 0.060 | 0.819 |
On CLINC150, out-of-scope recall is 0.59 (v0.3's 9B 0.46) with precision 0.84 (0.76).
Second temperature: v0.3.3 carries its own, 1.05, fitted by NLL on the MASSIVE items only (the two halves give 1.07 and 1.02). With the old 0.77 it would be overconfident (MASSIVE ECE 0.073, BANKING77 0.114).
Jev Decision Index and JevBench
JevK5-9B has not been run by the Jev Decision Index or submitted to JevBench yet. For reference, the index reran JevK5 0.2.2 (v0.2 4B weights, with the runtime that answers any number of options) on its previously refused rows (discussion #13). It scored 36.31, 15th of 49, and its calibration was 4th best (ECE 0.031). JevBench v1.4 ranked JevK5 v0.2 #2 of 76 systems.
How it was trained
On the same 47,460 rows as JevK5 v0.3 (4B); see that card for the full data table with licenses.
- Teacher questions (17,408): 3,270 from Qwen3.6-27B (Apache-2.0, self-hosted, thinking on) and 14,138 from GPT-6 Luna (OpenAI, through OpenAI's API). Each question was answered twice, independently, by its teacher and kept only when both answers matched the intended one. Luna's outputs were generated under OpenAI's terms, which govern their use; review them for your use case.
- Public replay (30,052 items): train splits of 26 public datasets, with no test or validation split of any dataset, and no split of MMLU, MMLU-Pro, ANLI or NLI4CT. Each dataset's license is in the JevK5 card's table.
- Training: cross-entropy on the option-letter logits, SemIf's prompt format, 1.5 epochs, learning rate 3e-5, inputs up to 2,048 tokens. One temperature (1.316) is fitted on the held-out teacher questions. The second temperature (1.05) is fitted by NLL on the 500 MASSIVE items above.
- Not released: the 2-epoch 9B (higher on teacher questions, 0.865, but lower on the index proxy
and the hard set; the pre-registered rule picked 1.5 epochs), and the two development 9Bs that
trained on ANLI/NLI4CT or on MMLU's
auxiliary_train(RACE, non-commercial), as described on the v0.3 card.
Data rules.
- No JevBench item, public or held out, and no output of Jev was used for training, tuning or selection. JevBench's public items were only used to report the numbers above.
- Every public and Luna row was checked against the test and validation text of 35 Decision Index benchmarks, and against JevBench's public items. A row was dropped for an exact match or for any shared 8-word sequence. GPQA and HLE are gated and were not checked; no source is built from them.
Declared overlap with the Jev Decision Index. These are train splits of index benchmarks, deduplicated against their test and validation items: ARC, OpenBookQA, CommonsenseQA, GSM8K, WinoGrande, HellaSwag, BANKING77, CLINC150, SGD, Amazon ESCI, When2Call, iSarcasmEval, RAGTruth, HoVer and the New Yorker caption contest. Separately, 53 of the Qwen-written training questions share at least one 8-word sequence with ContractNLI (41) or SGD (12) test or dev text; they were reported by the scan and kept.
Known weak spots
- One seed, small per-source dev slices. The sweep's winner rests on one run per setting, and each index source has about 40 dev rows.
- Calibration on the held-out teacher questions is a little worse than v0.3's 9B (ECE 0.043 against 0.032), and hard-tier calibration on JevBench is worse than the 4B's (0.071 against 0.054).
- On the dev set's teacher families, trap (0.76 → 0.64) and ambiguous (0.75 → 0.67) questions dropped against v0.3's 9B while judging (0.83 → 0.91) and dates and numbers (0.32 → 0.53) rose.
- Dates and numbers stay weak on JevBench's hard tier (0.47). On bev-decision, score questions are the weakest type (0.482), and BANKING77's calibration over 77 options is a little worse than v0.3's 9B (ECE 0.055 against 0.044).
- Accuracy drops sharply on JevBench's fresh sealed decisions (measured for JevK5 v0.2). Real-world workflow performance against Jev has not been measured.
- English only. Needs a CUDA GPU with ~19 GB for bf16. Inputs over 16,384 tokens are refused, not cut.
Use
from jevk5 import JevK5
model = JevK5("alibiserikbay/JevK5-9B")
model.decide(
"I was billed twice for order #4411. Please refund the duplicate charge today.",
{"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "tech": "Bugs", "sales": "New purchases"}},
)
# {'type': 'choice', 'confidence': 0.993, 'choice': 'billing', 'probabilities': {...}, ...}
With llama.cpp (see JevK5-GGUF):
from jevk5 import JevK5GGUF
model = JevK5GGUF("http://127.0.0.1:8080", temperature=1.316, knockout_temperature=1.05)
Or as a server that answers TypeSafe-style /v1/systemone requests:
jevk5-serve --model alibiserikbay/JevK5-9B --port 8090. Both read temperature and
knockout_temperature from this repo's jevk5_config.json.
Credits
Qwen3.5-9B and Qwen3.6-27B by the Qwen team (Apache-2.0). GPT-6 Luna by OpenAI. The one-pass readout and prompt come from SemIf by TheoLeeCJ (MIT). The public datasets listed above belong to their authors, under their licenses (table on the JevK5 card). Evaluated with JevBench (github.com/fstandhartinger/jevbench, MIT) and bev-decision-150K. Not affiliated with TypeSafe AI or Jev.
- Downloads last month
- 1,022