Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

AJev · Gemma 4 12B LoRA (lora5)

AJev is an open, Jev-style typed decision model. Give it a state (text or JSON) and typed questions — yes/no (noul), single choice (choice, up to 255 options) or ordered levels (score) — and it returns a calibrated probability for every option in a single forward pass per question. No text is generated.

This repository holds the LoRA adapter (r 32, α 64, 131M trainable parameters) for google/gemma-4-12B-it, with per-question-type calibration temperatures in ajev_lm_config.json. Inference code: github.com/cmzy/ajev-infer. The training code is not public.

Results

Decision Index 0.2.1 (the Jev Decision Index suite, complete run with the official kit and its scorer; no truncation, no option pruning; all 150,317 scoreable requests answered, HLE included):

Model Base Decision Index
Jev (TypeSafe, hosted) — 57.91
Surogate Rune 26B-A4B v3 (board #1) Gemma 4 26B 57.44
AJev lora5 (this model, self-run) Gemma 4 12B 52.22
reflex 27B Qwen 27B 52.16
Winnow-12B Gemma 4 12B 50.02
Jev-Omni Gemma 4 12B 40.53

Areas (chance-corrected): Knowledge & Reasoning 37.6, Language Understanding 57.9, Retrieval & Classification 56.1, Tools & Automation 68.8, Arts & Human Taste 37.1; raw (uncorrected) index 63.24. Strongest benchmarks: GSM8K 89.4, HellaSwag 90.3, BPoMP 85.7, When2Call 83.6, New Yorker 70.4 (chance-corrected skill, 0 = random, 100 = perfect). Weakest relative to the board: CRUXEval, VAST, API-Bank, HoVer, MMLU-Pro. The index above is our own run, not a maintainer reproduction.

Accuracy on held-out test sets (state up to 16K tokens, calibrated; zero-shot = the same base model and prompt):

Test set Size Zero-shot Gemma 4 12B AJev lora5
typed-decisions (business decisions) 2,000 0.696 0.791
JevBench public 231 0.848 0.853
Kev transfer v9 1,264 0.745 0.771
eikos-decisions heldout (en/es/pt) 1,190 0.928 0.931

Calibration (ECE) on typed-decisions drops from 0.280 (zero-shot) to 0.142; on Kev v9 from 0.204 to 0.060.

Usage

pip install "ajev-infer @ git+https://github.com/cmzy/ajev-infer"          # transformers
pip install "ajev-infer[vllm] @ git+https://github.com/cmzy/ajev-infer"    # + vLLM server
from ajev.lm.predictor import LMPredictor
from ajev.schema import decisions_from_jev, jev_answer

p = LMPredictor("google/gemma-4-12B-it", adapter="andyzhang232/ajev-gemma4-12b-lora5")
ds = decisions_from_jev(
    {"ticket": "I was charged twice for order #1182.", "tier": "gold"},
    {"topic": {"type": "choice", "instructions": "What is the ticket about?",
               "criteria": {"billing": "charges, refunds", "delivery": "shipping", "account": "login"}},
     "escalate": {"type": "noul", "instructions": "Should a human agent take this now?"}})
for d, probs in zip(ds, p.predict(ds)):
    print(d.meta["question_id"], jev_answer(d, probs))

Server with the Jev wire format (POST /v1/systemone, concurrent requests batched by vLLM):

python -m ajev.serve_vllm --adapter andyzhang232/ajev-gemma4-12b-lora5 --port 8000

Use transformers ≥ 5.17 (earlier versions tokenize Gemma 4 differently). Load the adapter as a LoRA (PEFT, or vLLM LoRARequest) or merge it in memory; a merged checkpoint saved with transformers 5.17 and reloaded gives different outputs.

How it works

Each question becomes one chat prompt: a fixed instruction line, the state, the question, a one-line hint for the question type, and the options labelled A, B, … (after Z, single-token two-letter codes such as AB, up to 255 options). The next-token logits of the labels (both the bare and the space-prefixed token) are combined, divided by a temperature fitted per question type on held-out data, and softmaxed. One forward pass per question; questions sharing a state can reuse its KV cache.

Training

  • Data (64,819 decisions, 1 epoch): public classification / NLI / preference / safety datasets in English and Chinese; bev-decision (long policies, traps, numeric, counterfactual); typed-decisions; programmatically generated business-rule cases (refund, invoice, contract, approval, SLA, tool use, chat routing, security logs, HR, triage); train splits of public benchmarks related to the leaderboard (HellaSwag, WinoGrande, GSM8K, MMLU auxiliary train, RAGTruth, ContractNLI, Humicroedit, NLI4CT, iSarcasm, ACOS, New Yorker, ANLI R1/R2, phishing); BANKING77/CLINC150 with all classes; When2Call preference data.
  • Decontamination: every training item was compared against the leaderboard's evaluation data (exact text match and shared ≥ 60-character sentences); 508 items were removed. Only train splits were used.
  • Objective: cross-entropy on the label logits (label smoothing 0.05), plus a ranked-probability term for score questions and a consistency term between option orders.
  • Setup: LoRA r 32 / α 64 on all attention and MLP projections of the language model; lr 3e-5, cosine; 32 decisions per step; states up to 16,384 tokens; one RTX PRO 6000 (96 GB), about 13 hours.

Limitations

  • JevBench's hard tier (0.703) and long-context items (> 1,500 tokens, 0.676) remain below the zero-shot base model; multi-hop lookup in long policies and date/number arithmetic are weak spots.
  • Knowledge-heavy benchmarks (MMLU-Pro, BBH, GPQA) are bounded by the 12B base.
  • Probabilities are calibrated on our own held-out mix; recalibrate on your own data when it differs a lot.
  • Several training datasets carry non-commercial or attribution terms. The adapter is released under Apache-2.0 for research use; check the terms of the underlying datasets before commercial deployment.
  • Not affiliated with TypeSafe AI.
Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyzhang232/ajev-gemma4-12b-lora5

Adapter
(106)
this model