Instructions to use frontier-infra/jebadiah-4b-v2-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use frontier-infra/jebadiah-4b-v2-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download frontier-infra/jebadiah-4b-v2-MLX --local-dir jebadiah-4b-v2-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Jebadiah 4B v2 MLX
MLX builds of Jebadiah 4B v2 for Apple silicon, one folder per precision (8bit/, 4bit/, group size 64).
Jebadiah answers a typed question (choice, noul or score) with a probability for every option, read from one
forward pass. Nothing is generated. Code, trainer and evals: getainode/jebadiah.
Results and docs
- Project site, with every result and how to run the models: jebadiah.ai.
- For the family: Decision Index 0.2.1: the 27B scores 54.67, #5 of 67 open models (as of 2026-09-26). This is the board's own number; the maintainer validated my run and put it on the leaderboard. Run record.
- JevBench v1.4.2: on its 231 public items, run through its own harness, Jebadiah 27B scores 0.866, the same as Jev 1.13.0; Jebadiah 9B v2 scores 0.818. This is my own run on the public items, not the official board, which also uses sealed items. Details and caveats. The 4B has not been run on either.
- All sizes: the Hugging Face collection, mirrored on ModelScope.
Which one should I use? For local use, start with Jebadiah 9B v2 GGUF. On Apple silicon, use an MLX build: 27B, 9B v2 or 4B v2. For vLLM or fine-tuning, use the full weights: 27B, 9B v2 or 4B v2. Setup for every runtime, and for JDE: Run Jeb locally.
Files
Every file was checked on the 260 held-out questions the merged weights were checked on, and compared with the merged bf16 weights and with the training run's own eval records.
| File | Size | Same answer as bf16 | Same as the run | choice + noul | score | Prob. diff median / max |
|---|---|---|---|---|---|---|
8bit/ |
4.5 GB | 252 / 260 | 253 / 260 | 169 / 173 | 84 / 87 | 0.004 / 0.110 |
4bit/ |
2.4 GB | 215 / 260 | 215 / 260 | 157 / 173 | 58 / 87 | 0.056 / 0.380 |
| bf16 weights | 259 / 260 | 172 / 173 | 87 / 87 | 0.002 / 0.019 |
Which one: 8bit if your Mac has the memory; 4bit when it does not. A folder needs about its own
size in unified memory, plus about 1 GB for a 2k-token prompt.
8bit changes 8 of 260 answers against bf16 (3 on score questions) and moves probabilities more (median 0.004, max 0.11). Use it only when a larger build does not fit.
4bit changes 45 of 260 answers against bf16 (29 on score questions) and moves probabilities more (median 0.056, max 0.38). Use it only when a larger build does not fit.
Run it
The script renders the prompt exactly as AINode does, runs one forward pass with mlx-lm, multiplies the
last hidden state by the option labels' output-head rows in fp32 and applies temperatures.json
(choice 1.1167, noul 1.3319, score 0.8312). Needs mlx-lm 0.31 or newer (qwen3_5 support).
pip install "mlx-lm>=0.31"
hf download frontier-infra/jebadiah-4b-v2-MLX --include "8bit/*" --include "scripts/*" --include temperatures.json --local-dir jebadiah-4b-v2-MLX
cd jebadiah-4b-v2-MLX
python scripts/decide_mlx.py --model 8bit --request scripts/example-request.json
--no-temperatures returns the raw probabilities.
HelpSteer2-like traffic. A held-out HelpSteer2 check (2026-09-29): on the 418 nvidia/HelpSteer2 validation rows that no reported evaluation set uses (natural label mix, never in the training pool), this model's score answers want a temperature of 1.26, not 0.83, and still 1.08 after reweighting to a flat label mix. If your traffic looks like HelpSteer2, keep the older score temperature, 1.20 (the train fit in temperatures.json; every script here takes --temperatures with a file of your own), or refit on your own labels. v3's calibration split will draw HelpSteer2 from held-out data.
On example-request.json (8bit/):
{
"route": {"type": "choice", "choice": "billing", "confidence": 0.394292, "probabilities": {"billing": 0.596195, "support": 0.0839, "sales": 0.319905}},
"urgent": {"type": "noul", "noul": 0.147991}
}
How it was measured
Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
20260925), the run's option order and the run's temperatures, so the probability differences compare like with like (the shipped temperatures.json since 2026-09-29 uses 0.8312 for score; top picks do not depend on it). "Same answer" is the top option; "prob. diff" is the
largest change on any option against the run's CUDA record. Records: eval/agreement-*.json.
License
Apache-2.0, as the base model. Made in Texas.
Quantized