Jebadiah 4B v2 MLX

MLX builds of Jebadiah 4B v2 for Apple silicon, one folder per precision (8bit/, 4bit/, group size 64). Jebadiah answers a typed question (choice, noul or score) with a probability for every option, read from one forward pass. Nothing is generated. Code, trainer and evals: getainode/jebadiah.

Results and docs

  • Project site, with every result and how to run the models: jebadiah.ai.
  • For the family: Decision Index 0.2.1: the 27B scores 54.67, #5 of 67 open models (as of 2026-09-26). This is the board's own number; the maintainer validated my run and put it on the leaderboard. Run record.
  • JevBench v1.4.2: on its 231 public items, run through its own harness, Jebadiah 27B scores 0.866, the same as Jev 1.13.0; Jebadiah 9B v2 scores 0.818. This is my own run on the public items, not the official board, which also uses sealed items. Details and caveats. The 4B has not been run on either.
  • All sizes: the Hugging Face collection, mirrored on ModelScope.

Which one should I use? For local use, start with Jebadiah 9B v2 GGUF. On Apple silicon, use an MLX build: 27B, 9B v2 or 4B v2. For vLLM or fine-tuning, use the full weights: 27B, 9B v2 or 4B v2. Setup for every runtime, and for JDE: Run Jeb locally.

Files

Every file was checked on the 260 held-out questions the merged weights were checked on, and compared with the merged bf16 weights and with the training run's own eval records.

File Size Same answer as bf16 Same as the run choice + noul score Prob. diff median / max
8bit/ 4.5 GB 252 / 260 253 / 260 169 / 173 84 / 87 0.004 / 0.110
4bit/ 2.4 GB 215 / 260 215 / 260 157 / 173 58 / 87 0.056 / 0.380
bf16 weights 259 / 260 172 / 173 87 / 87 0.002 / 0.019

Which one: 8bit if your Mac has the memory; 4bit when it does not. A folder needs about its own size in unified memory, plus about 1 GB for a 2k-token prompt.

8bit changes 8 of 260 answers against bf16 (3 on score questions) and moves probabilities more (median 0.004, max 0.11). Use it only when a larger build does not fit. 4bit changes 45 of 260 answers against bf16 (29 on score questions) and moves probabilities more (median 0.056, max 0.38). Use it only when a larger build does not fit.

Run it

The script renders the prompt exactly as AINode does, runs one forward pass with mlx-lm, multiplies the last hidden state by the option labels' output-head rows in fp32 and applies temperatures.json (choice 1.1167, noul 1.3319, score 0.8312). Needs mlx-lm 0.31 or newer (qwen3_5 support).

pip install "mlx-lm>=0.31"
hf download frontier-infra/jebadiah-4b-v2-MLX --include "8bit/*" --include "scripts/*" --include temperatures.json --local-dir jebadiah-4b-v2-MLX
cd jebadiah-4b-v2-MLX
python scripts/decide_mlx.py --model 8bit --request scripts/example-request.json

--no-temperatures returns the raw probabilities.

HelpSteer2-like traffic. A held-out HelpSteer2 check (2026-09-29): on the 418 nvidia/HelpSteer2 validation rows that no reported evaluation set uses (natural label mix, never in the training pool), this model's score answers want a temperature of 1.26, not 0.83, and still 1.08 after reweighting to a flat label mix. If your traffic looks like HelpSteer2, keep the older score temperature, 1.20 (the train fit in temperatures.json; every script here takes --temperatures with a file of your own), or refit on your own labels. v3's calibration split will draw HelpSteer2 from held-out data.

On example-request.json (8bit/):

{
 "route": {"type": "choice", "choice": "billing", "confidence": 0.394292, "probabilities": {"billing": 0.596195, "support": 0.0839, "sales": 0.319905}},
 "urgent": {"type": "noul", "noul": 0.147991}
}

How it was measured

Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed 20260925), the run's option order and the run's temperatures, so the probability differences compare like with like (the shipped temperatures.json since 2026-09-29 uses 0.8312 for score; top picks do not depend on it). "Same answer" is the top option; "prob. diff" is the largest change on any option against the run's CUDA record. Records: eval/agreement-*.json.

License

Apache-2.0, as the base model. Made in Texas.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for frontier-infra/jebadiah-4b-v2-MLX

Finetuned
Qwen/Qwen3.5-4B
Quantized
(3)
this model

Collection including frontier-infra/jebadiah-4b-v2-MLX