Instructions to use andyzhang232/ajev-gemma4-12b-lora5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use andyzhang232/ajev-gemma4-12b-lora5 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
AJev · Gemma 4 12B LoRA (lora5)
AJev is an open, Jev-style typed decision model. Give it a state (text or JSON) and typed questions —
yes/no (noul), single choice (choice, up to 255 options) or ordered levels (score) — and it returns a
calibrated probability for every option in a single forward pass per question. No text is generated.
This repository holds the LoRA adapter (r 32, α 64, 131M trainable parameters) for google/gemma-4-12B-it,
with per-question-type calibration temperatures in ajev_lm_config.json.
Inference code: github.com/cmzy/ajev-infer. The training code is not public.
Results
Decision Index 0.2.1 (the Jev Decision Index suite, complete run with the official kit and its scorer; no truncation, no option pruning; all 150,317 scoreable requests answered, HLE included):
| Model | Base | Decision Index |
|---|---|---|
| Jev (TypeSafe, hosted) | — | 57.91 |
| Surogate Rune 26B-A4B v3 (board #1) | Gemma 4 26B | 57.44 |
| AJev lora5 (this model, self-run) | Gemma 4 12B | 52.22 |
| reflex 27B | Qwen 27B | 52.16 |
| Winnow-12B | Gemma 4 12B | 50.02 |
| Jev-Omni | Gemma 4 12B | 40.53 |
Areas (chance-corrected): Knowledge & Reasoning 37.6, Language Understanding 57.9, Retrieval & Classification 56.1, Tools & Automation 68.8, Arts & Human Taste 37.1; raw (uncorrected) index 63.24. Strongest benchmarks: GSM8K 89.4, HellaSwag 90.3, BPoMP 85.7, When2Call 83.6, New Yorker 70.4 (chance-corrected skill, 0 = random, 100 = perfect). Weakest relative to the board: CRUXEval, VAST, API-Bank, HoVer, MMLU-Pro. The index above is our own run, not a maintainer reproduction.
Accuracy on held-out test sets (state up to 16K tokens, calibrated; zero-shot = the same base model and prompt):
| Test set | Size | Zero-shot Gemma 4 12B | AJev lora5 |
|---|---|---|---|
| typed-decisions (business decisions) | 2,000 | 0.696 | 0.791 |
| JevBench public | 231 | 0.848 | 0.853 |
| Kev transfer v9 | 1,264 | 0.745 | 0.771 |
| eikos-decisions heldout (en/es/pt) | 1,190 | 0.928 | 0.931 |
Calibration (ECE) on typed-decisions drops from 0.280 (zero-shot) to 0.142; on Kev v9 from 0.204 to 0.060.
Usage
pip install "ajev-infer @ git+https://github.com/cmzy/ajev-infer" # transformers
pip install "ajev-infer[vllm] @ git+https://github.com/cmzy/ajev-infer" # + vLLM server
from ajev.lm.predictor import LMPredictor
from ajev.schema import decisions_from_jev, jev_answer
p = LMPredictor("google/gemma-4-12B-it", adapter="andyzhang232/ajev-gemma4-12b-lora5")
ds = decisions_from_jev(
{"ticket": "I was charged twice for order #1182.", "tier": "gold"},
{"topic": {"type": "choice", "instructions": "What is the ticket about?",
"criteria": {"billing": "charges, refunds", "delivery": "shipping", "account": "login"}},
"escalate": {"type": "noul", "instructions": "Should a human agent take this now?"}})
for d, probs in zip(ds, p.predict(ds)):
print(d.meta["question_id"], jev_answer(d, probs))
Server with the Jev wire format (POST /v1/systemone, concurrent requests batched by vLLM):
python -m ajev.serve_vllm --adapter andyzhang232/ajev-gemma4-12b-lora5 --port 8000
Use transformers ≥ 5.17 (earlier versions tokenize Gemma 4 differently). Load the adapter as a LoRA (PEFT, or vLLM
LoRARequest) or merge it in memory; a merged checkpoint saved with transformers 5.17 and reloaded gives
different outputs.
How it works
Each question becomes one chat prompt: a fixed instruction line, the state, the question, a one-line hint for the
question type, and the options labelled A, B, … (after Z, single-token two-letter codes such as AB, up to 255
options). The next-token logits of the labels (both the bare and the space-prefixed token) are combined, divided by
a temperature fitted per question type on held-out data, and softmaxed. One forward pass per question; questions
sharing a state can reuse its KV cache.
Training
- Data (64,819 decisions, 1 epoch): public classification / NLI / preference / safety datasets in English and Chinese; bev-decision (long policies, traps, numeric, counterfactual); typed-decisions; programmatically generated business-rule cases (refund, invoice, contract, approval, SLA, tool use, chat routing, security logs, HR, triage); train splits of public benchmarks related to the leaderboard (HellaSwag, WinoGrande, GSM8K, MMLU auxiliary train, RAGTruth, ContractNLI, Humicroedit, NLI4CT, iSarcasm, ACOS, New Yorker, ANLI R1/R2, phishing); BANKING77/CLINC150 with all classes; When2Call preference data.
- Decontamination: every training item was compared against the leaderboard's evaluation data (exact text match and shared ≥ 60-character sentences); 508 items were removed. Only train splits were used.
- Objective: cross-entropy on the label logits (label smoothing 0.05), plus a ranked-probability term for
scorequestions and a consistency term between option orders. - Setup: LoRA r 32 / α 64 on all attention and MLP projections of the language model; lr 3e-5, cosine; 32 decisions per step; states up to 16,384 tokens; one RTX PRO 6000 (96 GB), about 13 hours.
Limitations
- JevBench's hard tier (0.703) and long-context items (> 1,500 tokens, 0.676) remain below the zero-shot base model; multi-hop lookup in long policies and date/number arithmetic are weak spots.
- Knowledge-heavy benchmarks (MMLU-Pro, BBH, GPQA) are bounded by the 12B base.
- Probabilities are calibrated on our own held-out mix; recalibrate on your own data when it differs a lot.
- Several training datasets carry non-commercial or attribution terms. The adapter is released under Apache-2.0 for research use; check the terms of the underlying datasets before commercial deployment.
- Not affiliated with TypeSafe AI.
- Downloads last month
- 20