cloudyu commited on
Commit
dcc8186
·
verified ·
1 Parent(s): bd5c55d

Add files using upload-large-folder tool

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,314 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ license_link: https://ai.google.dev/gemma/docs/gemma_4_license
4
+ base_model: google/gemma-4-26B-A4B-it
5
+ base_model_relation: adapter
6
+ language:
7
+ - en
8
+ library_name: transformers
9
+ pipeline_tag: text-classification
10
+ tags:
11
+ - system-one
12
+ - system-two
13
+ - adaptive-thinking
14
+ - typed-decisions
15
+ - decision-model
16
+ - calibrated-probabilities
17
+ - jev
18
+ - noul
19
+ - choice
20
+ - score
21
+ - lora
22
+ - gemma4
23
+ - mixture-of-experts
24
+ - multimodal
25
+ - vllm
26
+ model-index:
27
+ - name: autotrust/GEV-26B-Decide
28
+ results:
29
+ - task:
30
+ type: text-classification
31
+ name: Jev Decision Index 0.2.1, System 1 (complete run, 150,317 scored requests)
32
+ dataset:
33
+ type: decision-index
34
+ name: Decision Index suite 0.2 (edition 0.2.1)
35
+ metrics:
36
+ - type: decision_index
37
+ name: Decision Index (balanced skill)
38
+ value: 58.05
39
+ ---
40
+
41
+ # autotrust/GEV-26B-Decide
42
+
43
+ ### A decision model with adaptive thinking, on Gemma-4-26B-A4B-it (26 B parameters, ≈ 4 B active per token)
44
+
45
+ **GEV-26B-Decide answers typed questions with a calibrated probability for every option, and thinks only when it needs
46
+ to.** System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same
47
+ backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities.
48
+ One set of weights, one vLLM engine, for text and images.
49
+
50
+ | | what it does | output |
51
+ |---|---|---|
52
+ | **System 1** | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
53
+ | **Adaptive thinking** | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
54
+ | **System 2** | the unmodified `google/gemma-4-26B-A4B-it`, optionally thinking step by step, text and images | text / reasoning |
55
+
56
+ GEV-26B-Decide was previously published as `autotrust/JEV-Gemma4-26B-A4B`; the weights are the same.
57
+
58
+ > **Two models, two organisations.** **TypeSafe Jev 1.13** is the hosted, closed model made by TypeSafe AI.
59
+ > **autotrust/GEV-26B-Decide** is an independent open-weights model built by AutoTrust AI; it is not affiliated with,
60
+ > endorsed by, or a product of TypeSafe AI.
61
+
62
+ ## Adaptive thinking
63
+
64
+ 1. **Fast distribution.** System 1 returns p1 in one pass.
65
+ 2. **Think only when uncertain.** If the leading option of p1 is below the threshold (default 0.8), System 2 reasons in
66
+ Gemma-4's thinking mode over the same state, question and options (budget 8,192 thinking tokens).
67
+ 3. **Fold the reasoning in.** When the thinking channel closes, the answer-letter distribution p2 is read in one step,
68
+ and the result is p = ½ p1 + ½ p2. On its own, p2 is close to one-hot and over-confident; the equal mix keeps the
69
+ reasoning's accuracy and System 1's calibration.
70
+
71
+ The threshold and the mix were chosen on 1,754 questions from six public sets that are not part of the Decision Index
72
+ (test or validation splits, 300 random questions each; AQuA-RAT has 254):
73
+
74
+ | set | System 1 | **adaptive** | thinking on | always think |
75
+ |---|---:|---:|---:|---:|
76
+ | AQuA-RAT (math word problems) | 68.1 | **89.0** | 53.9 % | 90.2 |
77
+ | LogiQA (logical reasoning) | 55.3 | **80.3** | 63.7 % | 82.0 |
78
+ | StrategyQA (multi-hop yes/no) | 68.0 | **78.7** | 71.0 % | 79.0 |
79
+ | MedMCQA (medical) | 66.7 | **71.7** | 52.3 % | 73.7 |
80
+ | OpenBookQA (science) | 94.3 | **96.3** | 14.7 % | 95.7 |
81
+ | CommonsenseQA | 86.7 | **85.0** | 32.0 % | 83.3 |
82
+ | **all 1,754** | 73.3 | **83.4** | 47.8 % | 83.8 |
83
+
84
+ Accuracy in %. Calibration is unchanged: ECE 0.035 for System 1 and 0.035 for adaptive. The adaptive mode reaches 96 % of
85
+ the always-think gain while thinking on 48 % of the questions. Thinking length: median 2,282 tokens, 90th percentile 7,573;
86
+ 91.6 % of the thoughts finish within the 8,192-token budget.
87
+
88
+ **Decision Index, Knowledge & Reasoning with adaptive thinking:** the full run of the area's ten benchmarks (33,846
89
+ requests) is in progress; its results will be added here.
90
+
91
+ ## Decision Index 0.2.1 (System 1)
92
+
93
+ Complete run of all 150,759 requests of suite 0.2 (150,317 scored; 0 errors, 0 unsupported) with System 1 only, scored
94
+ with the kit's `score --edition 0.2.1`. Results:
95
+ [`autotrust/jev-decision-index-results`](https://huggingface.co/datasets/autotrust/jev-decision-index-results)
96
+ (`runs/jev-gemma4-26b-a4b`, the run of these weights under their previous name).
97
+
98
+ | | Decision Index (balanced skill) | balanced raw | breadth skill |
99
+ |---|---:|---:|---:|
100
+ | **autotrust/GEV-26B-Decide** (System 1) | **58.05** | 67.35 | 56.98 |
101
+ | TypeSafe Jev 1.13 (board) | 57.91 | — | — |
102
+ | [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B) | 53.30 | 64.32 | 52.21 |
103
+
104
+ | area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
105
+ |---|---:|---:|---:|---:|---:|
106
+ | GEV-26B-Decide (System 1) | 0.430 | 0.636 | 0.679 | 0.697 | 0.415 |
107
+
108
+ ## Images
109
+
110
+ The checkpoint contains Gemma-4's vision encoder (no audio encoder), so System 1 and System 2 both accept images. The
111
+ decision head was trained on text; decisions over images are zero-shot.
112
+
113
+ | check | result |
114
+ |---|---|
115
+ | synthetic images: colour (8 options), shape (4), printed number (8), "is there a red object?" (yes/no) | 100 % on each (30 images each) |
116
+ | [VL-RewardBench](https://huggingface.co/datasets/MMInstruction/VL-RewardBench), 1,247 pairs, both presentation orders averaged | **78.4 %** overall (general 55.8, hallucination 84.9, reasoning 76.0; macro 72.2) |
117
+
118
+ For reference, [autotrust/JEV-27B-VL](https://huggingface.co/autotrust/JEV-27B-VL) scores 78.3 % on VL-RewardBench with
119
+ the same protocol.
120
+
121
+ ## Context length
122
+
123
+ The backbone's native context is **262,144 tokens (256K)**. We tested decisions that hinge on a single sentence placed at
124
+ a random depth in long real text (concatenated PubMedQA abstracts): a yes/no question and a 16-option question, 10 of
125
+ each per length, on one B200 with vLLM, one request at a time.
126
+
127
+ | prompt length | yes/no correct | 16-option correct | mean probability on the right answer | median latency |
128
+ |---|---:|---:|---:|---:|
129
+ | 4K | 10/10 | 10/10 | 0.999 | 0.15 s |
130
+ | 32K | 10/10 | 10/10 | 0.999 | 1.6 s |
131
+ | 64K | 10/10 | 10/10 | 0.999 | 5.0 s |
132
+ | 128K | 10/10 | 10/10 | 1.000 | 17.8 s |
133
+
134
+ Lengths above 128K have not been tested yet.
135
+
136
+ ## Many options
137
+
138
+ `choice` takes 2–256 options. Up to 16 are read in one pass with the trained labels A–P. More options are read in groups
139
+ of at most 16 (in parallel), then a final of 16; every option is read and none is pruned (`strategy: "tournament"`, the
140
+ default). Zero-shot intent classification, all options offered at once, 400 test utterances per row:
141
+
142
+ | test | options | accuracy |
143
+ |---|---:|---:|
144
+ | MASSIVE (en) | 59 | 91.2 % |
145
+ | BANKING77 | 77 | 81.5 % |
146
+ | CLINC150 | 150 | 95.5 % |
147
+ | CLINC150 utterances among CLINC150 + BANKING77 + MASSIVE intents | 255 | 89.0 % |
148
+ | BANKING77 utterances among the same 255 intents | 255 | 75.2 % |
149
+
150
+ The 255-option sets merge three catalogues with overlapping intents, so part of the drop comes from near-duplicate labels.
151
+ A single pass with labels beyond P (`strategy: "single"`) is about 3× faster but less accurate here (CLINC150: 88.5 %
152
+ against 95.2 %), so it is not the default. The BANKING77 and CLINC150 training splits are part of the training data (see
153
+ below).
154
+
155
+ ## Quick start (vLLM)
156
+
157
+ ```bash
158
+ hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide
159
+ bash GEV-26B-Decide/serve.sh # vLLM on :8000; one GPU with 80 GB or more
160
+ ```
161
+
162
+ `serve.sh` runs `serve_decide.py`: the standard vLLM OpenAI server (same flags as `vllm serve`) with a `POST /v1/decide`
163
+ route. It loads the backbone once: plain requests are System 2, and requests for the LoRA module `jev-decision`
164
+ (`adapter_vllm/`: backbone LoRA + the decision head as an `lm_head` LoRA) are System 1. It needs a vLLM build with
165
+ Gemma-4 support plus `patches/vllm-gemma4-lm-head-lora.patch` (LoRA on Gemma-4's tied `lm_head`, vocabulary 262,144);
166
+ tested with a vLLM development build from September 2026.
167
+
168
+ ### System 1 and adaptive thinking: `POST /v1/decide`
169
+
170
+ ```bash
171
+ curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
172
+ "kind": "choice",
173
+ "state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
174
+ "question": "How much does the ball cost?",
175
+ "options": ["$0.10", "$0.05", "$1.00", "$0.55"],
176
+ "thinking": "auto"}'
177
+ ```
178
+
179
+ | field | value |
180
+ |---|---|
181
+ | `kind` | `noul`: yes/no, probabilities for `["false", "true"]` · `score`: 0–5 · `choice`: your `options` |
182
+ | `state` | what the decision is about: a string, a JSON object, or a list mixing text and images `["Photo: ", {"image": "https://… or data:…"}]` |
183
+ | `question` | one question about the state |
184
+ | `options` | `choice` only: 2–256 strings |
185
+ | `thinking` | `"off"` (default: System 1 only), `"auto"` (adaptive), `"on"` (always think); `noul` and `choice` |
186
+ | `threshold` | System 1 confidence below which `"auto"` thinks (default 0.8) |
187
+ | `think_budget` | maximum thinking tokens (default 8,192) |
188
+ | `strategy` | more than 16 options: `"tournament"` (default), `"single"`, `"permute"` |
189
+ | `return_reasoning` / `debug` | include System 2's reasoning / the System 1 and System 2 distributions |
190
+
191
+ The response has `options`, `probabilities`, `choice`, `choice_index`, `usage` and, when thinking was requested,
192
+ `thinking: {"used": true, "think_tokens": …, "think_seconds": …}`. `GET /v1/decide/info` lists the defaults.
193
+
194
+ ```python
195
+ import requests
196
+
197
+ def decide(kind, state, question, options=None, thinking="auto"):
198
+ body = {"kind": kind, "state": state, "question": question, "thinking": thinking, **({"options": options} if options else {})}
199
+ r = requests.post("http://localhost:8000/v1/decide", json=body).json()
200
+ return dict(zip(r["options"], r["probabilities"])), r.get("thinking", {}).get("used")
201
+
202
+ decide("noul", "John was born on 29 February 1996.", "Was John's 7th birthday celebrated on a 29 February?")
203
+ ```
204
+
205
+ ### System 2
206
+
207
+ ```python
208
+ requests.post("http://localhost:8000/v1/chat/completions", json={
209
+ "model": "autotrust/GEV-26B-Decide",
210
+ "messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
211
+ "max_tokens": 200, "chat_template_kwargs": {"enable_thinking": False}})
212
+ ```
213
+
214
+ ### Speed (one B200, vLLM)
215
+
216
+ * System 1: median 45 ms for a single request; 257 decisions per second with 64 concurrent clients. vLLM matches the
217
+ `transformers` engine below to a mean largest probability difference of 0.015 (300 held-out decisions).
218
+ * Adaptive thinking adds a few seconds when it thinks (median 2,282 thinking tokens) and nothing when it does not.
219
+
220
+ If you call `/v1/completions` for System 1 yourself, pass `top_k: 0` and `top_p: 1.0`: the model's generation config sets
221
+ `top_k=64` and `top_p=0.95`, which vLLM applies as request defaults and which would truncate the returned probabilities.
222
+
223
+ ## Usage (transformers + peft, System 1)
224
+
225
+ ```python
226
+ import json, torch
227
+ from huggingface_hub import snapshot_download
228
+ from peft import PeftModel
229
+ from safetensors.torch import load_file
230
+ from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
231
+
232
+ d = snapshot_download("autotrust/GEV-26B-Decide")
233
+ tok = AutoTokenizer.from_pretrained(d)
234
+ base = Gemma4ForConditionalGeneration.from_pretrained(d, dtype=torch.bfloat16, device_map="cuda")
235
+ # System 2: `base` is gemma-4-26B-A4B-it unchanged; use base.generate(...) (text or images).
236
+
237
+ # System 1: adapter merged in memory + head
238
+ m = PeftModel.from_pretrained(base, f"{d}/adapter").merge_and_unload().eval()
239
+ backbone = m.model
240
+ jc, T = json.load(open(f"{d}/judge_config.json")), json.load(open(f"{d}/calibration.json"))["per_kind"]
241
+ head = load_file(f"{d}/head.safetensors"); W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
242
+
243
+ @torch.no_grad()
244
+ def decide(kind, state, question, options):
245
+ lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
246
+ text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
247
+ ids = torch.tensor([[tok.bos_token_id] + tok.encode(text, add_special_tokens=False)], device="cuda")
248
+ h = backbone(input_ids=ids, use_cache=False).last_hidden_state[0, -1].float()
249
+ z = 30.0 * torch.tanh((W @ h + b) / 30.0)
250
+ s, _ = jc["slots"]["ranges"][kind]
251
+ return dict(zip(options, torch.softmax(z[s:s + len(options)] / T[kind], 0).tolist()))
252
+
253
+ print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
254
+ "Is the customer asking for a refund?", ["false", "true"]))
255
+ ```
256
+
257
+ This path reads up to 16 options per pass; for more, use the server (or read groups of 16 and a final, as above).
258
+
259
+ ## Engine and read-out
260
+
261
+ * **Template** `bare-v1`, prefixed with `<bos>`: `[kind] … [state] … [question] … [options] A) … [decision]:`
262
+ * **Read-out**: the final-norm hidden state of the last token goes through a linear fp32 head (hidden 2,816 → 24 slots),
263
+ soft-capped at 30 like Gemma's own logits; inactive slots are masked, the logits are divided by the per-kind
264
+ temperature and softmaxed. Nothing is generated.
265
+
266
+ ## Temperatures
267
+
268
+ | table | noul | choice | score |
269
+ |---|---|---|---|
270
+ | `calibration.json` (default; used in the Decision Index run) | 1.003 | 1.017 | 0.999 |
271
+ | `calibration_gold.json` (calibrated against ground-truth answers) | 1.214 | 1.098 | 1.000 |
272
+
273
+ Use `calibration_gold.json` when you gate automatic actions on confidence.
274
+
275
+ ## Training data (disclosure)
276
+
277
+ System 1 was trained on teacher distributions and ground-truth decision data. The ground-truth data includes the public
278
+ **training splits** of some datasets whose **test** splits the Decision Index uses (among them BANKING77 and CLINC150);
279
+ no test split of any benchmark was used, and suite items were excluded before training. The list has been provided to
280
+ the Decision Index maintainers. MMMU / MMMU-Pro are not valid evaluations for this model. The adaptive-thinking settings
281
+ were chosen on data outside the Decision Index suite.
282
+
283
+ ## Limitations
284
+
285
+ * Where the teacher is wrong, System 1 often is too; adaptive thinking helps most on math, logic and multi-step
286
+ questions, and can slightly lower accuracy on commonsense questions (CommonsenseQA 86.7 → 85.0).
287
+ * Thinking is slow on hard inputs: on expert-level and chess questions it often uses the full 8,192-token budget.
288
+ * Weaker than JEV-27B on long structured inputs (e.g. the Decision Index's Home appliance simulator and POP909).
289
+ * HLE: below chance with System 1, like every open entry on the board.
290
+ * Decisions over images are zero-shot; contexts above 128K tokens are untested.
291
+ * The one-engine vLLM setup needs the bundled vLLM patch; `noul` and `score` accept only their canonical options.
292
+ * English-centric; not for high-stakes decisions without confidence gating.
293
+
294
+ ## Files
295
+
296
+ ```
297
+ model-*.safetensors · config.json · processor_config.json · tokenizer* · chat_template.jinja · generation_config.json
298
+ google/gemma-4-26B-A4B-it, unchanged (System 2; text + image input)
299
+ adapter/ System 1 LoRA (peft), for the transformers path
300
+ head.safetensors 24-slot decision head (fp32): proj.weight [24, 2816], proj.bias [24]
301
+ judge_config.json slot layout, verbalizer ids, softcap, read-out
302
+ calibration.json per-kind temperatures (default)
303
+ calibration_gold.json per-kind temperatures calibrated against ground-truth answers
304
+ adapter_vllm/ System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA, plus decision_head.json
305
+ serve_decide.py · serve.sh
306
+ vLLM server with POST /v1/decide (System 1, adaptive thinking) next to the OpenAI endpoints
307
+ patches/ vLLM patch: LoRA on Gemma-4's tied lm_head
308
+ reports/ Decision Index scores (System 1); adaptive-thinking validation summary
309
+ ```
310
+
311
+ ## License
312
+
313
+ Apache-2.0 for the adapter, head and calibration files; base model under the Gemma 4 terms
314
+ (<https://ai.google.dev/gemma/docs/gemma_4_license>). Not affiliated with TypeSafe AI.
adapter/README.md ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ # System 1 adapter (PEFT LoRA)
2
+
3
+ LoRA adapter for [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it), part of
4
+ [autotrust/GEV-26B-Decide](https://huggingface.co/autotrust/GEV-26B-Decide). Use it with `head.safetensors`, `judge_config.json`
5
+ and `calibration.json` from the repository root; see the model card. `adapter_vllm/` is the same System 1 for vLLM.
adapter/adapter_config.json ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": {
6
+ "base_model_class": "Gemma4ForConditionalGeneration",
7
+ "parent_library": "transformers.models.gemma4.modeling_gemma4"
8
+ },
9
+ "base_model_name_or_path": "google/gemma-4-26B-A4B-it",
10
+ "bias": "none",
11
+ "corda_config": null,
12
+ "ensure_weight_tying": false,
13
+ "eva_config": null,
14
+ "exclude_modules": null,
15
+ "fan_in_fan_out": false,
16
+ "inference_mode": true,
17
+ "init_lora_weights": true,
18
+ "kasa_config": null,
19
+ "layer_replication": null,
20
+ "layers_pattern": null,
21
+ "layers_to_transform": null,
22
+ "loftq_config": {},
23
+ "lora_alpha": 64,
24
+ "lora_bias": false,
25
+ "lora_dropout": 0.05,
26
+ "lora_ga_config": null,
27
+ "megatron_config": null,
28
+ "megatron_core": "megatron.core",
29
+ "modules_to_save": null,
30
+ "monteclora_config": null,
31
+ "peft_type": "LORA",
32
+ "peft_version": "0.21.0",
33
+ "qalora_group_size": 16,
34
+ "r": 32,
35
+ "rank_pattern": {},
36
+ "revision": null,
37
+ "target_modules": "model\\.language_model\\.layers\\.\\d+\\.(self_attn\\.(q|k|v|o)_proj|mlp\\.(gate|up|down)_proj)",
38
+ "target_parameters": null,
39
+ "task_type": null,
40
+ "trainable_token_indices": null,
41
+ "use_bdlora": null,
42
+ "use_dora": false,
43
+ "use_qalora": false,
44
+ "use_rslora": false,
45
+ "velora_config": null
46
+ }
adapter/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:94fb34ad3b58266c304a58d1675ae4a9f02f3be453d08135e185c8079236836e
3
+ size 148745744
adapter_vllm/adapter_config.json ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": {
6
+ "base_model_class": "Gemma4ForConditionalGeneration",
7
+ "parent_library": "transformers.models.gemma4.modeling_gemma4"
8
+ },
9
+ "base_model_name_or_path": "google/gemma-4-26B-A4B-it",
10
+ "bias": "none",
11
+ "corda_config": null,
12
+ "ensure_weight_tying": false,
13
+ "eva_config": null,
14
+ "exclude_modules": null,
15
+ "fan_in_fan_out": false,
16
+ "inference_mode": true,
17
+ "init_lora_weights": true,
18
+ "kasa_config": null,
19
+ "layer_replication": null,
20
+ "layers_pattern": null,
21
+ "layers_to_transform": null,
22
+ "loftq_config": {},
23
+ "lora_alpha": 64.0,
24
+ "lora_bias": false,
25
+ "lora_dropout": 0.0,
26
+ "lora_ga_config": null,
27
+ "megatron_config": null,
28
+ "megatron_core": "megatron.core",
29
+ "modules_to_save": null,
30
+ "monteclora_config": null,
31
+ "peft_type": "LORA",
32
+ "peft_version": "0.21.0",
33
+ "qalora_group_size": 16,
34
+ "r": 32,
35
+ "rank_pattern": {},
36
+ "revision": null,
37
+ "target_modules": [
38
+ "down_proj",
39
+ "gate_proj",
40
+ "k_proj",
41
+ "lm_head",
42
+ "o_proj",
43
+ "q_proj",
44
+ "up_proj",
45
+ "v_proj"
46
+ ],
47
+ "target_parameters": null,
48
+ "task_type": null,
49
+ "trainable_token_indices": null,
50
+ "use_bdlora": null,
51
+ "use_dora": false,
52
+ "use_qalora": false,
53
+ "use_rslora": false,
54
+ "velora_config": null
55
+ }
adapter_vllm/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6c20767b1a3fd8717d5cf709a98276666d2c6d7e84e973026a07091343294291
3
+ size 92591488
adapter_vllm/decision_head.json ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bias": [
3
+ 0.004050370771437883,
4
+ 0.0025548420380800962,
5
+ -0.0009425673633813858,
6
+ -0.00034990705898962915,
7
+ 0.0007876436575315893,
8
+ 0.004409449640661478,
9
+ 0.002710630651563406,
10
+ -0.0009728983859531581,
11
+ -0.0007952140877023339,
12
+ -0.0021821260452270508,
13
+ -0.001166780013591051,
14
+ -0.00140042370185256,
15
+ -0.0026794816367328167,
16
+ -0.003741167951375246,
17
+ -0.003188050352036953,
18
+ -0.006115781608968973,
19
+ -0.0021734880283474922,
20
+ -0.0018658898770809174,
21
+ -0.0028179585933685303,
22
+ -0.00424833782017231,
23
+ -0.004550046753138304,
24
+ -0.004217465873807669,
25
+ -0.005448763258755207,
26
+ -0.005735703278332949
27
+ ],
28
+ "verbalizer_ids": [
29
+ 4530,
30
+ 3397,
31
+ 236771,
32
+ 236770,
33
+ 236778,
34
+ 236800,
35
+ 236812,
36
+ 236810,
37
+ 236776,
38
+ 236799,
39
+ 236780,
40
+ 236796,
41
+ 236788,
42
+ 236811,
43
+ 236823,
44
+ 236814,
45
+ 236777,
46
+ 236863,
47
+ 236855,
48
+ 236798,
49
+ 236792,
50
+ 236797,
51
+ 236806,
52
+ 236791
53
+ ],
54
+ "slots": {
55
+ "num_slots": 24,
56
+ "ranges": {
57
+ "noul": [
58
+ 0,
59
+ 2
60
+ ],
61
+ "score": [
62
+ 2,
63
+ 8
64
+ ],
65
+ "choice": [
66
+ 8,
67
+ 24
68
+ ]
69
+ },
70
+ "verbalizers": [
71
+ "false",
72
+ "true",
73
+ "0",
74
+ "1",
75
+ "2",
76
+ "3",
77
+ "4",
78
+ "5",
79
+ "A",
80
+ "B",
81
+ "C",
82
+ "D",
83
+ "E",
84
+ "F",
85
+ "G",
86
+ "H",
87
+ "I",
88
+ "J",
89
+ "K",
90
+ "L",
91
+ "M",
92
+ "N",
93
+ "O",
94
+ "P"
95
+ ],
96
+ "template_version": "bare-v1"
97
+ },
98
+ "softcap": 30.0,
99
+ "prefix_token": 2,
100
+ "note": "decision logits = lm_head-LoRA logprobs of verbalizer ids + bias, then per-kind temperature (Gemma: vLLM applies the model's final-logit softcap; |bias| is tiny, so adding it after the cap is exact to ~1e-2)"
101
+ }
calibration.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "version": 1,
3
+ "per_kind": {
4
+ "noul": 1.0027465035537586,
5
+ "choice": 1.017196267473986,
6
+ "score": 0.9992906948414192
7
+ },
8
+ "per_kind_family": {},
9
+ "fit": {
10
+ "noul": {
11
+ "n": 4363,
12
+ "T": 1.0027465035537586,
13
+ "kl_before": 0.007547731976956129,
14
+ "kl_after": 0.007547072134912014
15
+ },
16
+ "choice": {
17
+ "n": 3418,
18
+ "T": 1.017196267473986,
19
+ "kl_before": 0.06508523225784302,
20
+ "kl_after": 0.06501992046833038
21
+ },
22
+ "score": {
23
+ "n": 3173,
24
+ "T": 0.9992906948414192,
25
+ "kl_before": 0.035448525100946426,
26
+ "kl_after": 0.035448361188173294
27
+ }
28
+ },
29
+ "diagnostic_calibration_split": {
30
+ "raw": {
31
+ "kl": 0.03358318656682968,
32
+ "ece": 0.001969979859321292,
33
+ "mce": 0.01303844153881073
34
+ },
35
+ "calibrated": {
36
+ "kl": 0.033562496304512024,
37
+ "ece": 0.0012317931094873945,
38
+ "mce": 0.012980207800865173
39
+ }
40
+ },
41
+ "source": "checkpoints/s2_gemma4_step7000",
42
+ "fit_rows": 10954,
43
+ "d1_excluded": true
44
+ }
calibration_gold.json ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "version": 1,
3
+ "per_kind": {
4
+ "noul": 1.2136210312191784,
5
+ "choice": 1.0979579333985496,
6
+ "score": 1.0
7
+ },
8
+ "per_kind_family": {},
9
+ "fit": {
10
+ "noul": {
11
+ "n": 763,
12
+ "T": 1.2136210312191784,
13
+ "kl_before": 0.3925037682056427,
14
+ "kl_after": 0.38618776202201843
15
+ },
16
+ "choice": {
17
+ "n": 17929,
18
+ "T": 1.0979579333985496,
19
+ "kl_before": 0.318155437707901,
20
+ "kl_after": 0.31654489040374756
21
+ }
22
+ },
23
+ "diagnostic_calibration_split": {
24
+ "raw": {
25
+ "kl": 0.3211902976036072,
26
+ "ece": 0.013034004604066436,
27
+ "mce": 0.0645633339881897
28
+ },
29
+ "calibrated": {
30
+ "kl": 0.31938767433166504,
31
+ "ece": 0.005258639238451115,
32
+ "mce": 0.8051199913024902
33
+ }
34
+ },
35
+ "source": "checkpoints/s2_gemma4_step7000",
36
+ "fit_rows": 18692,
37
+ "d1_excluded": true
38
+ }
chat_template.jinja ADDED
@@ -0,0 +1,390 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {#
2
+ Template: Google Gemma 4 Canonical Chat Template
3
+ Author: Google Gemma Engineering Team
4
+ Published: 2026-07-09
5
+ Context: Fixed tool-calling loops, turn closures, and thinking content-ordering.
6
+ #}
7
+ {%- macro format_parameters(properties, required, filter_keys=false) -%}
8
+ {%- set standard_keys = ['description', 'type', 'properties', 'required', 'nullable'] -%}
9
+ {%- set ns = namespace(found_first=false) -%}
10
+ {%- for key, value in properties | dictsort -%}
11
+ {%- set add_comma = false -%}
12
+ {%- if not filter_keys or key not in standard_keys -%}
13
+ {%- if ns.found_first %},{% endif -%}
14
+ {%- set ns.found_first = true -%}
15
+ {{ key }}:{
16
+ {%- if value['description'] -%}
17
+ description:<|"|>{{ value['description'] }}<|"|>
18
+ {%- set add_comma = true -%}
19
+ {%- endif -%}
20
+ {%- if value['type'] | upper == 'STRING' -%}
21
+ {%- if value['enum'] -%}
22
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
23
+ enum:{{ format_argument(value['enum']) }}
24
+ {%- endif -%}
25
+ {%- elif value['type'] | upper == 'ARRAY' -%}
26
+ {%- if value['items'] is mapping and value['items'] -%}
27
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
28
+ items:{
29
+ {%- set ns_items = namespace(found_first=false) -%}
30
+ {%- for item_key, item_value in value['items'] | dictsort -%}
31
+ {%- if item_value is not none -%}
32
+ {%- if ns_items.found_first %},{% endif -%}
33
+ {%- set ns_items.found_first = true -%}
34
+ {%- if item_key == 'properties' -%}
35
+ properties:{
36
+ {%- if item_value is mapping -%}
37
+ {{- format_parameters(item_value, value['items']['required'] | default([])) -}}
38
+ {%- endif -%}
39
+ }
40
+ {%- elif item_key == 'required' -%}
41
+ required:[
42
+ {%- for req_item in item_value -%}
43
+ <|"|>{{- req_item -}}<|"|>
44
+ {%- if not loop.last %},{% endif -%}
45
+ {%- endfor -%}
46
+ ]
47
+ {%- elif item_key == 'type' -%}
48
+ {%- if item_value is string -%}
49
+ type:{{ format_argument(item_value | upper) }}
50
+ {%- else -%}
51
+ type:{{ format_argument(item_value | map('upper') | list) }}
52
+ {%- endif -%}
53
+ {%- else -%}
54
+ {{ item_key }}:{{ format_argument(item_value) }}
55
+ {%- endif -%}
56
+ {%- endif -%}
57
+ {%- endfor -%}
58
+ }
59
+ {%- endif -%}
60
+ {%- endif -%}
61
+ {%- if value['nullable'] %}
62
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
63
+ nullable:true
64
+ {%- endif -%}
65
+ {%- if value['type'] | upper == 'OBJECT' -%}
66
+ {%- if value['properties'] is defined and value['properties'] is mapping -%}
67
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
68
+ properties:{
69
+ {{- format_parameters(value['properties'], value['required'] | default([])) -}}
70
+ }
71
+ {%- elif value is mapping -%}
72
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
73
+ properties:{
74
+ {{- format_parameters(value, value['required'] | default([]), filter_keys=true) -}}
75
+ }
76
+ {%- endif -%}
77
+ {%- if value['required'] -%}
78
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
79
+ required:[
80
+ {%- for item in value['required'] | default([]) -%}
81
+ <|"|>{{- item -}}<|"|>
82
+ {%- if not loop.last %},{% endif -%}
83
+ {%- endfor -%}
84
+ ]
85
+ {%- endif -%}
86
+ {%- endif -%}
87
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
88
+ type:<|"|>{{ value['type'] | upper }}<|"|>}
89
+ {%- endif -%}
90
+ {%- endfor -%}
91
+ {%- endmacro -%}
92
+ {%- macro format_function_declaration(tool_data) -%}
93
+ declaration:{{- tool_data['function']['name'] -}}{description:<|"|>{{- tool_data['function']['description'] -}}<|"|>
94
+ {%- set params = tool_data['function']['parameters'] -%}
95
+ {%- if params -%}
96
+ ,parameters:{
97
+ {%- if params['properties'] -%}
98
+ properties:{ {{- format_parameters(params['properties'], params['required']) -}} },
99
+ {%- endif -%}
100
+ {%- if params['required'] -%}
101
+ required:[
102
+ {%- for item in params['required'] -%}
103
+ <|"|>{{- item -}}<|"|>
104
+ {{- ',' if not loop.last -}}
105
+ {%- endfor -%}
106
+ ],
107
+ {%- endif -%}
108
+ {%- if params['type'] -%}
109
+ type:<|"|>{{- params['type'] | upper -}}<|"|>}
110
+ {%- endif -%}
111
+ {%- endif -%}
112
+ {%- if 'response' in tool_data['function'] -%}
113
+ {%- set response_declaration = tool_data['function']['response'] -%}
114
+ ,response:{
115
+ {%- if response_declaration['description'] -%}
116
+ description:<|"|>{{- response_declaration['description'] -}}<|"|>,
117
+ {%- endif -%}
118
+ {%- if response_declaration['type'] | upper == 'OBJECT' -%}
119
+ type:<|"|>{{- response_declaration['type'] | upper -}}<|"|>}
120
+ {%- endif -%}
121
+ {%- endif -%}
122
+ }
123
+ {%- endmacro -%}
124
+ {%- macro format_argument(argument, escape_keys=True) -%}
125
+ {%- if argument is none -%}
126
+ {{- 'null' -}}
127
+ {%- elif argument is string -%}
128
+ {{- '<|"|>' + argument + '<|"|>' -}}
129
+ {%- elif argument is boolean -%}
130
+ {{- 'true' if argument else 'false' -}}
131
+ {%- elif argument is mapping -%}
132
+ {{- '{' -}}
133
+ {%- set ns = namespace(found_first=false) -%}
134
+ {%- for key, value in argument | dictsort -%}
135
+ {%- if ns.found_first %},{% endif -%}
136
+ {%- set ns.found_first = true -%}
137
+ {%- if escape_keys -%}
138
+ {{- '<|"|>' + key + '<|"|>' -}}
139
+ {%- else -%}
140
+ {{- key -}}
141
+ {%- endif -%}
142
+ :{{- format_argument(value, escape_keys=escape_keys) -}}
143
+ {%- endfor -%}
144
+ {{- '}' -}}
145
+ {%- elif argument is sequence -%}
146
+ {{- '[' -}}
147
+ {%- for item in argument -%}
148
+ {{- format_argument(item, escape_keys=escape_keys) -}}
149
+ {%- if not loop.last %},{% endif -%}
150
+ {%- endfor -%}
151
+ {{- ']' -}}
152
+ {%- else -%}
153
+ {{- argument -}}
154
+ {%- endif -%}
155
+ {%- endmacro -%}
156
+ {%- macro strip_thinking(text) -%}
157
+ {%- set ns = namespace(result='') -%}
158
+ {%- for part in text.split('<channel|>') -%}
159
+ {%- if '<|channel>' in part -%}
160
+ {%- set ns.result = ns.result + part.split('<|channel>')[0] -%}
161
+ {%- else -%}
162
+ {%- set ns.result = ns.result + part -%}
163
+ {%- endif -%}
164
+ {%- endfor -%}
165
+ {{- ns.result | trim -}}
166
+ {%- endmacro -%}
167
+
168
+ {%- macro format_tool_response_block(tool_name, response) -%}
169
+ {{- '<|tool_response>' -}}
170
+ {%- if response is mapping -%}
171
+ {{- 'response:' + tool_name + '{' -}}
172
+ {%- for key, value in response | dictsort -%}
173
+ {{- key -}}:{{- format_argument(value, escape_keys=False) -}}
174
+ {%- if not loop.last %},{% endif -%}
175
+ {%- endfor -%}
176
+ {{- '}' -}}
177
+ {%- else -%}
178
+ {{- 'response:' + tool_name + '{value:' + format_argument(response, escape_keys=False) + '}' -}}
179
+ {%- endif -%}
180
+ {{- '<tool_response|>' -}}
181
+ {%- endmacro -%}
182
+
183
+ {#- ===== SETUP ===== -#}
184
+ {%- set ns = namespace(prev_message_type=None, prev_non_tool_role=None) -%}
185
+ {%- set loop_messages = messages -%}
186
+ {%- set enable_thinking = enable_thinking | default(false) -%}
187
+ {%- set preserve_thinking = preserve_thinking | default(false) -%}
188
+ {{- bos_token -}}
189
+ {#- Handle System/Tool Definitions Block -#}
190
+ {%- if enable_thinking or tools or (messages and messages[0]['role'] in ['system', 'developer']) -%}
191
+ {{- '<|turn>system\n' -}}
192
+ {#- Inject Thinking token at the very top of the FIRST system turn -#}
193
+ {%- if enable_thinking -%}
194
+ {{- '<|think|>\n' -}}
195
+ {%- set ns.prev_message_type = 'think' -%}
196
+ {%- endif -%}
197
+ {%- if messages and messages[0]['role'] in ['system', 'developer'] -%}
198
+ {%- if messages[0]['content'] is string -%}
199
+ {{- messages[0]['content'] | trim -}}
200
+ {%- elif messages[0]['content'] is sequence -%}
201
+ {%- for item in messages[0]['content'] -%}
202
+ {{- item['text'] | trim + ' '-}}
203
+ {%- endfor -%}
204
+ {%- endif -%}
205
+ {%- set loop_messages = messages[1:] -%}
206
+ {%- endif -%}
207
+ {%- if tools -%}
208
+ {%- for tool in tools %}
209
+ {{- '<|tool>' -}}
210
+ {{- format_function_declaration(tool) | trim -}}
211
+ {{- '<tool|>' -}}
212
+ {%- endfor %}
213
+ {%- set ns.prev_message_type = 'tool' -%}
214
+ {%- endif -%}
215
+ {{- '<turn|>\n' -}}
216
+ {%- endif %}
217
+
218
+ {#- Pre-scan: find last user message index for reasoning guard -#}
219
+ {%- set ns_turn = namespace(last_user_idx=-1) -%}
220
+ {%- for i in range(loop_messages | length) -%}
221
+ {%- if loop_messages[i]['role'] == 'user' -%}
222
+ {%- set ns_turn.last_user_idx = i -%}
223
+ {%- endif -%}
224
+ {%- endfor -%}
225
+
226
+ {#- Loop through messages -#}
227
+ {%- for message in loop_messages -%}
228
+ {%- if message['role'] != 'tool' -%}
229
+ {%- set ns.prev_message_type = None -%}
230
+ {%- set role = 'model' if message['role'] == 'assistant' else message['role'] -%}
231
+ {#- Detect continuation using tracked state — O(1) instead of O(n) backward scan -#}
232
+ {%- set continue_same_model_turn = (role == 'model' and ns.prev_non_tool_role == 'assistant') -%}
233
+ {%- if not continue_same_model_turn -%}
234
+ {{- '<|turn>' + role + '\n' }}
235
+
236
+ {%- endif -%}
237
+
238
+ {#- Render reasoning/reasoning_content as thinking channel -#}
239
+ {%- set thinking_text = message.get('reasoning') or message.get('reasoning_content') -%}
240
+ {%- set thinking_gate = (loop.index0 > ns_turn.last_user_idx) or (preserve_thinking and message.get('tool_calls')) -%}
241
+ {%- if thinking_text and thinking_gate -%}
242
+ {{- '<|channel>thought\n' + thinking_text + '\n<channel|>' -}}
243
+ {%- endif -%}
244
+
245
+ {%- if message.get('tool_calls') -%}
246
+ {%- for tool_call in message.get('tool_calls') -%}
247
+ {%- set function = tool_call['function'] -%}
248
+ {{- '<|tool_call>call:' + function['name'] + '{' -}}
249
+ {%- if function['arguments'] is mapping -%}
250
+ {%- set ns_args = namespace(found_first=false) -%}
251
+ {%- for key, value in function['arguments'] | dictsort -%}
252
+ {%- if ns_args.found_first %},{% endif -%}
253
+ {%- set ns_args.found_first = true -%}
254
+ {{- key -}}:{{- format_argument(value, escape_keys=False) -}}
255
+ {%- endfor -%}
256
+ {%- elif function['arguments'] is none -%}
257
+ {%- else -%}
258
+ {{- raise_exception(
259
+ "chat_template: tool_calls[].function.arguments must be a "
260
+ "JSON object (mapping), not a string. Deserialize arguments "
261
+ "before passing to the template."
262
+ ) -}}
263
+ {%- endif -%}
264
+ {{- '}<tool_call|>' -}}
265
+ {%- endfor -%}
266
+ {%- set ns.prev_message_type = 'tool_call' -%}
267
+ {%- endif -%}
268
+
269
+ {%- set ns_tr_out = namespace(flag=false) -%}
270
+ {%- if message.get('tool_responses') -%}
271
+ {#- Legacy: tool_responses embedded on the assistant message (Google/Gemma native) -#}
272
+ {%- for tool_response in message.get('tool_responses') -%}
273
+ {{- format_tool_response_block(tool_response['name'] | default('unknown', true), tool_response['response']) -}}
274
+ {%- set ns_tr_out.flag = true -%}
275
+ {%- set ns.prev_message_type = 'tool_response' -%}
276
+ {%- endfor -%}
277
+ {%- elif message.get('tool_calls') -%}
278
+ {#- OpenAI Chat Completions: forward-scan consecutive role:tool messages -#}
279
+ {%- set ns_tool_scan = namespace(stopped=false) -%}
280
+ {%- for k in range(loop.index0 + 1, loop_messages | length) -%}
281
+ {%- if ns_tool_scan.stopped -%}
282
+ {%- elif loop_messages[k]['role'] != 'tool' -%}
283
+ {%- set ns_tool_scan.stopped = true -%}
284
+ {%- else -%}
285
+ {%- set follow = loop_messages[k] -%}
286
+ {#- Resolve tool_call_id to function name -#}
287
+ {%- set ns_tname = namespace(name=follow.get('name') or 'unknown') -%}
288
+ {%- for tc in message.get('tool_calls') -%}
289
+ {%- if tc.get('id') == follow.get('tool_call_id') -%}
290
+ {%- set ns_tname.name = tc['function']['name'] -%}
291
+ {%- endif -%}
292
+ {%- endfor -%}
293
+ {#- Handle content as string or content-parts array -#}
294
+ {%- set tool_body = follow.get('content') -%}
295
+ {%- if tool_body is string -%}
296
+ {{- format_tool_response_block(ns_tname.name, tool_body) -}}
297
+ {%- elif tool_body is sequence and tool_body is not string -%}
298
+ {%- set ns_txt = namespace(s='') -%}
299
+ {%- for part in tool_body -%}
300
+ {%- if part.get('type') == 'text' -%}
301
+ {%- set ns_txt.s = ns_txt.s + (part.get('text') | default('')) -%}
302
+ {%- endif -%}
303
+ {%- endfor -%}
304
+ {{- format_tool_response_block(ns_tname.name, ns_txt.s) -}}
305
+ {%- for part in tool_body -%}
306
+ {%- if part.get('type') in ['image', 'image_url'] -%}
307
+ {{- '<|image|>' -}}
308
+ {%- elif part.get('type') in ['audio', 'input_audio'] -%}
309
+ {{- '<|audio|>' -}}
310
+ {%- elif part.get('type') == 'video' -%}
311
+ {{- '<|video|>' -}}
312
+ {%- endif -%}
313
+ {%- endfor -%}
314
+ {%- else -%}
315
+ {{- format_tool_response_block(ns_tname.name, tool_body) -}}
316
+ {%- endif -%}
317
+ {%- set ns_tr_out.flag = true -%}
318
+ {%- set ns.prev_message_type = 'tool_response' -%}
319
+ {%- endif -%}
320
+ {%- endfor -%}
321
+ {%- endif -%}
322
+
323
+ {%- set captured_content -%}
324
+ {%- if message.get('content') is string -%}
325
+ {%- if role == 'model' -%}
326
+ {{- strip_thinking(message['content']) -}}
327
+ {%- else -%}
328
+ {{- message['content'] | trim -}}
329
+ {%- endif -%}
330
+ {%- elif message.get('content') is sequence -%}
331
+ {%- for item in message['content'] -%}
332
+ {%- if item.get('type') == 'text' -%}
333
+ {%- if role == 'model' -%}
334
+ {{- strip_thinking(item['text']) -}}
335
+ {%- else -%}
336
+ {{- item['text'] | trim -}}
337
+ {%- endif -%}
338
+ {%- elif item.get('type') in ['image', 'image_url'] -%}
339
+ {{- '<|image|>' -}}
340
+ {%- elif item.get('type') in ['audio', 'input_audio'] -%}
341
+ {{- '<|audio|>' -}}
342
+ {%- elif item.get('type') == 'video' -%}
343
+ {{- '<|video|>' -}}
344
+ {%- endif -%}
345
+ {%- endfor -%}
346
+ {%- endif -%}
347
+ {%- endset -%}
348
+
349
+ {{- captured_content -}}
350
+ {%- set has_content = captured_content | trim | length > 0 -%}
351
+
352
+ {#- Forward-scan: find next non-tool message role for continuation detection -#}
353
+ {%- set next_nt = namespace(role=None, found=false) -%}
354
+ {%- for j in range(loop.index0 + 1, loop_messages | length) -%}
355
+ {%- if not next_nt.found -%}
356
+ {%- if loop_messages[j]['role'] != 'tool' -%}
357
+ {%- set next_nt.role = loop_messages[j]['role'] -%}
358
+ {%- set next_nt.found = true -%}
359
+ {%- endif -%}
360
+ {%- endif -%}
361
+ {%- endfor -%}
362
+
363
+ {%- set continues_into_next = (
364
+ role == 'model'
365
+ and next_nt.role == 'assistant'
366
+ and (not message.get('tool_calls') or ns_tr_out.flag)
367
+ ) -%}
368
+
369
+ {%- if ns.prev_message_type == 'tool_call' and not ns_tr_out.flag -%}
370
+ {{- '<|tool_response>' -}}
371
+ {%- elif continues_into_next -%}
372
+ {%- elif not (ns_tr_out.flag and not has_content and not next_nt.found) -%}
373
+ {{- '<turn|>\n' -}}
374
+ {%- endif -%}
375
+
376
+ {#- Track previous non-tool role for next iteration (avoids O(n) backward scan) -#}
377
+ {%- set ns.prev_non_tool_role = message['role'] -%}
378
+ {%- endif -%}
379
+ {%- endfor -%}
380
+
381
+ {%- if add_generation_prompt -%}
382
+ {%- if ns.prev_message_type != 'tool_response' and ns.prev_message_type != 'tool_call' -%}
383
+ {{- '<|turn>model\n' -}}
384
+ {%- if not enable_thinking -%}
385
+ {{- '<|channel>thought\n<channel|>' -}}
386
+ {%- endif -%}
387
+ {%- elif ns.prev_message_type == 'tool_response' and enable_thinking -%}
388
+ {{- '<|channel>thought\n' -}}
389
+ {%- endif -%}
390
+ {%- endif -%}
config.json ADDED
@@ -0,0 +1,146 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Gemma4ForConditionalGeneration"
4
+ ],
5
+ "audio_config": null,
6
+ "audio_token_id": 258881,
7
+ "boa_token_id": 256000,
8
+ "boi_token_id": 255999,
9
+ "dtype": "bfloat16",
10
+ "eoa_token_id": 258883,
11
+ "eoa_token_index": 258883,
12
+ "eoi_token_id": 258882,
13
+ "eos_token_id": [
14
+ 1,
15
+ 106
16
+ ],
17
+ "image_token_id": 258880,
18
+ "initializer_range": 0.02,
19
+ "model_type": "gemma4",
20
+ "text_config": {
21
+ "attention_bias": false,
22
+ "attention_dropout": 0.0,
23
+ "attention_k_eq_v": true,
24
+ "bos_token_id": 2,
25
+ "dtype": "bfloat16",
26
+ "enable_moe_block": true,
27
+ "eos_token_id": 1,
28
+ "final_logit_softcapping": 30.0,
29
+ "global_head_dim": 512,
30
+ "head_dim": 256,
31
+ "hidden_activation": "gelu_pytorch_tanh",
32
+ "hidden_size": 2816,
33
+ "hidden_size_per_layer_input": 0,
34
+ "initializer_range": 0.02,
35
+ "intermediate_size": 2112,
36
+ "layer_types": [
37
+ "sliding_attention",
38
+ "sliding_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "sliding_attention",
42
+ "full_attention",
43
+ "sliding_attention",
44
+ "sliding_attention",
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "sliding_attention",
48
+ "full_attention",
49
+ "sliding_attention",
50
+ "sliding_attention",
51
+ "sliding_attention",
52
+ "sliding_attention",
53
+ "sliding_attention",
54
+ "full_attention",
55
+ "sliding_attention",
56
+ "sliding_attention",
57
+ "sliding_attention",
58
+ "sliding_attention",
59
+ "sliding_attention",
60
+ "full_attention",
61
+ "sliding_attention",
62
+ "sliding_attention",
63
+ "sliding_attention",
64
+ "sliding_attention",
65
+ "sliding_attention",
66
+ "full_attention"
67
+ ],
68
+ "max_position_embeddings": 262144,
69
+ "model_type": "gemma4_text",
70
+ "moe_intermediate_size": 704,
71
+ "num_attention_heads": 16,
72
+ "num_experts": 128,
73
+ "num_global_key_value_heads": 2,
74
+ "num_hidden_layers": 30,
75
+ "num_key_value_heads": 8,
76
+ "num_kv_shared_layers": 0,
77
+ "pad_token_id": 0,
78
+ "rms_norm_eps": 1e-06,
79
+ "rope_parameters": {
80
+ "full_attention": {
81
+ "partial_rotary_factor": 0.25,
82
+ "rope_theta": 1000000.0,
83
+ "rope_type": "proportional"
84
+ },
85
+ "sliding_attention": {
86
+ "rope_theta": 10000.0,
87
+ "rope_type": "default"
88
+ }
89
+ },
90
+ "sliding_window": 1024,
91
+ "tie_word_embeddings": true,
92
+ "top_k_experts": 8,
93
+ "use_bidirectional_attention": "vision",
94
+ "use_cache": true,
95
+ "use_double_wide_mlp": false,
96
+ "vocab_size": 262144,
97
+ "vocab_size_per_layer_input": 262144
98
+ },
99
+ "tie_word_embeddings": true,
100
+ "transformers_version": "5.5.0.dev0",
101
+ "video_token_id": 258884,
102
+ "vision_config": {
103
+ "_name_or_path": "",
104
+ "architectures": null,
105
+ "attention_bias": false,
106
+ "attention_dropout": 0.0,
107
+ "chunk_size_feed_forward": 0,
108
+ "default_output_length": 280,
109
+ "dtype": "bfloat16",
110
+ "global_head_dim": 72,
111
+ "head_dim": 72,
112
+ "hidden_activation": "gelu_pytorch_tanh",
113
+ "hidden_size": 1152,
114
+ "id2label": {
115
+ "0": "LABEL_0",
116
+ "1": "LABEL_1"
117
+ },
118
+ "initializer_range": 0.02,
119
+ "intermediate_size": 4304,
120
+ "is_encoder_decoder": false,
121
+ "label2id": {
122
+ "LABEL_0": 0,
123
+ "LABEL_1": 1
124
+ },
125
+ "max_position_embeddings": 131072,
126
+ "model_type": "gemma4_vision",
127
+ "num_attention_heads": 16,
128
+ "num_hidden_layers": 27,
129
+ "num_key_value_heads": 16,
130
+ "output_attentions": false,
131
+ "output_hidden_states": false,
132
+ "patch_size": 16,
133
+ "pooling_kernel_size": 3,
134
+ "position_embedding_size": 10240,
135
+ "problem_type": null,
136
+ "return_dict": true,
137
+ "rms_norm_eps": 1e-06,
138
+ "rope_parameters": {
139
+ "rope_theta": 100.0,
140
+ "rope_type": "default"
141
+ },
142
+ "standardize": true,
143
+ "use_clipped_linears": false
144
+ },
145
+ "vision_soft_tokens_per_image": 280
146
+ }
generation_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 2,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 1,
6
+ 106,
7
+ 50
8
+ ],
9
+ "pad_token_id": 0,
10
+ "temperature": 1.0,
11
+ "top_k": 64,
12
+ "top_p": 0.95,
13
+ "transformers_version": "5.5.0.dev0"
14
+ }
head.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8ffb37579b962dfb52c3b9684446649635b5a78ca3f7f4210c889421bf5851ad
3
+ size 270584
judge_config.json ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_model_path": ".",
3
+ "hidden_size": 2816,
4
+ "slots": {
5
+ "num_slots": 24,
6
+ "ranges": {
7
+ "noul": [
8
+ 0,
9
+ 2
10
+ ],
11
+ "score": [
12
+ 2,
13
+ 8
14
+ ],
15
+ "choice": [
16
+ 8,
17
+ 24
18
+ ]
19
+ },
20
+ "verbalizers": [
21
+ "false",
22
+ "true",
23
+ "0",
24
+ "1",
25
+ "2",
26
+ "3",
27
+ "4",
28
+ "5",
29
+ "A",
30
+ "B",
31
+ "C",
32
+ "D",
33
+ "E",
34
+ "F",
35
+ "G",
36
+ "H",
37
+ "I",
38
+ "J",
39
+ "K",
40
+ "L",
41
+ "M",
42
+ "N",
43
+ "O",
44
+ "P"
45
+ ],
46
+ "template_version": "bare-v1"
47
+ },
48
+ "verbalizer_ids": [
49
+ 4530,
50
+ 3397,
51
+ 236771,
52
+ 236770,
53
+ 236778,
54
+ 236800,
55
+ 236812,
56
+ 236810,
57
+ 236776,
58
+ 236799,
59
+ 236780,
60
+ 236796,
61
+ 236788,
62
+ 236811,
63
+ 236823,
64
+ 236814,
65
+ 236777,
66
+ 236863,
67
+ 236855,
68
+ 236798,
69
+ 236792,
70
+ 236797,
71
+ 236806,
72
+ 236791
73
+ ],
74
+ "kinds": [
75
+ "noul",
76
+ "choice",
77
+ "score"
78
+ ],
79
+ "softcap": 30.0,
80
+ "readout": {
81
+ "type": "last_token",
82
+ "canvas_len": 1,
83
+ "canvas_token": 0,
84
+ "prefix_token": 2,
85
+ "no_pad_mask": true
86
+ },
87
+ "stage": "s2",
88
+ "lora": {
89
+ "r": 32,
90
+ "alpha": 64,
91
+ "dropout": 0.05,
92
+ "target_modules": "model\\.language_model\\.layers\\.\\d+\\.(self_attn\\.(q|k|v|o)_proj|mlp\\.(gate|up|down)_proj)"
93
+ },
94
+ "model_name": "gev-26b-a4b",
95
+ "model_version": "0.1.0",
96
+ "adapter_subfolder": "adapter",
97
+ "weights_mode": "dual",
98
+ "base_model": "google/gemma-4-26B-A4B-it",
99
+ "temperatures": "jev"
100
+ }
model-00001-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1127684971bbca40465435a5cad69d67ad603bf5e61c6dfd5561fae4a3bcfdb3
3
+ size 49907246508
model-00002-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aab47033e1e8a492ef8e581efae1cf36478d0433567e7729b3c1728bc8970db7
3
+ size 1704763408
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
patches/vllm-gemma4-lm-head-lora.patch ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ From f8a48c40b339da4bdfeb86cf8bf179f5e4e397d9 Mon Sep 17 00:00:00 2001
2
+ From: yuhai-china <5901109+yuhai-china@users.noreply.github.com>
3
+ Date: Tue, 29 Sep 2026 16:22:28 +0000
4
+ Subject: [PATCH 1/2] [LoRA] Gemma-4: allow lm_head LoRA (declare
5
+ embedding_modules; tied lm_head) and raise LoRA vocab limit to 262144
6
+
7
+ ---
8
+ vllm/lora/layers/logits_processor.py | 6 ++++--
9
+ vllm/model_executor/models/gemma4.py | 7 +++++++
10
+ 2 files changed, 11 insertions(+), 2 deletions(-)
11
+
12
+ diff --git a/vllm/lora/layers/logits_processor.py b/vllm/lora/layers/logits_processor.py
13
+ index 0199437..33b6a2d 100644
14
+ --- a/vllm/lora/layers/logits_processor.py
15
+ +++ b/vllm/lora/layers/logits_processor.py
16
+ @@ -88,8 +88,10 @@ class LogitsProcessorWithLoRA(BaseLayerWithLoRA):
17
+ model_config: PretrainedConfig | None = None,
18
+ ) -> None:
19
+ # TODO: Verify if this condition can be further relaxed
20
+ - if self.base_layer.vocab_size > 258048:
21
+ - raise ValueError("When using LoRA, vocab size must be <= 258048")
22
+ + # Raised from 258048 to 262144 so that Gemma-4 (vocab 262,144) can carry an
23
+ + # lm_head LoRA; verified by output parity for Gemma-4-26B-A4B.
24
+ + if self.base_layer.vocab_size > 262144:
25
+ + raise ValueError("When using LoRA, vocab size must be <= 262144")
26
+ self.lora_a_stacked = torch.zeros(
27
+ (
28
+ max_loras,
29
+ diff --git a/vllm/model_executor/models/gemma4.py b/vllm/model_executor/models/gemma4.py
30
+ index 7f850f4..95c29de 100644
31
+ --- a/vllm/model_executor/models/gemma4.py
32
+ +++ b/vllm/model_executor/models/gemma4.py
33
+ @@ -1548,6 +1548,13 @@ class Gemma4ForCausalLM(
34
+ "up_proj",
35
+ ],
36
+ }
37
+ + # Maps PEFT lm_head LoRA targets onto vLLM's logits-processor LoRA wrapper.
38
+ + # Gemma-4 ties lm_head to the input embeddings; the LoRA delta is applied to
39
+ + # the output logits only (before the final-logit soft cap), so the input
40
+ + # embeddings and the base model's logits are unchanged when no LoRA is active.
41
+ + embedding_modules = {
42
+ + "lm_head": "output_embeddings",
43
+ + }
44
+
45
+ def __init__(self, *, vllm_config: VllmConfig, prefix: str = ""):
46
+ config = _get_text_config(vllm_config.model_config.hf_config)
47
+ --
48
+ 2.34.1
49
+
50
+
51
+ From c7083ba5922c4822fe5a1efb7e9bd0d71c905ec1 Mon Sep 17 00:00:00 2001
52
+ From: yuhai-china <5901109+yuhai-china@users.noreply.github.com>
53
+ Date: Tue, 29 Sep 2026 21:46:44 +0000
54
+ Subject: [PATCH 2/2] [LoRA] Dedupe aliased modules at activation (backport of
55
+ vllm-project/vllm#39816; fixes Gemma-4 LoRA silently reset)
56
+
57
+ ---
58
+ vllm/lora/model_manager.py | 13 +++++++++++++
59
+ 1 file changed, 13 insertions(+)
60
+
61
+ diff --git a/vllm/lora/model_manager.py b/vllm/lora/model_manager.py
62
+ index d19bf31..8804dcb 100644
63
+ --- a/vllm/lora/model_manager.py
64
+ +++ b/vllm/lora/model_manager.py
65
+ @@ -336,8 +336,21 @@ class LoRAModelManager:
66
+ "Activating LoRA. int id: %d, slot index: %d", lora_model.id, index
67
+ )
68
+ self.lora_index_to_id[index] = lora_model.id
69
+ + # Aliased module paths (e.g. Gemma-4 `model.layers.N` and
70
+ + # `model.self_decoder.decoder_layers.N`) point to the same physical LoRA
71
+ + # wrapper: pick one entry per physical module, preferring the path that has
72
+ + # weights, so a later alias cannot reset weights that were just loaded
73
+ + # (upstream vllm-project/vllm#39816, fixes #39815).
74
+ + chosen_modules: dict = {}
75
+ for module_name, module in self.modules.items():
76
+ + module_key = id(module)
77
+ module_lora = self._get_lora_layer_weights(lora_model, module_name)
78
+ + if module_key not in chosen_modules:
79
+ + chosen_modules[module_key] = (module_name, module, module_lora)
80
+ + elif module_lora is not None and chosen_modules[module_key][2] is None:
81
+ + chosen_modules[module_key] = (module_name, module, module_lora)
82
+ +
83
+ + for module_name, module, module_lora in chosen_modules.values():
84
+ if not module_lora:
85
+ module.reset_lora(index)
86
+ logger.debug(
87
+ --
88
+ 2.34.1
89
+
processor_config.json ADDED
@@ -0,0 +1,75 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "audio_ms_per_token": 40,
3
+ "audio_seq_length": 750,
4
+ "feature_extractor": {
5
+ "dither": 0.0,
6
+ "feature_extractor_type": "Gemma4AudioFeatureExtractor",
7
+ "feature_size": 128,
8
+ "fft_length": 512,
9
+ "fft_overdrive": false,
10
+ "frame_length": 320,
11
+ "hop_length": 160,
12
+ "input_scale_factor": 1.0,
13
+ "max_frequency": 8000.0,
14
+ "mel_floor": 0.001,
15
+ "min_frequency": 0.0,
16
+ "padding_side": "right",
17
+ "padding_value": 0.0,
18
+ "per_bin_mean": null,
19
+ "per_bin_stddev": null,
20
+ "preemphasis": 0.0,
21
+ "preemphasis_htk_flavor": true,
22
+ "return_attention_mask": true,
23
+ "sampling_rate": 16000
24
+ },
25
+ "image_processor": {
26
+ "do_convert_rgb": true,
27
+ "do_normalize": false,
28
+ "do_rescale": true,
29
+ "do_resize": true,
30
+ "image_mean": [
31
+ 0.0,
32
+ 0.0,
33
+ 0.0
34
+ ],
35
+ "image_processor_type": "Gemma4ImageProcessor",
36
+ "image_seq_length": 280,
37
+ "image_std": [
38
+ 1.0,
39
+ 1.0,
40
+ 1.0
41
+ ],
42
+ "max_soft_tokens": 280,
43
+ "patch_size": 16,
44
+ "pooling_kernel_size": 3,
45
+ "resample": 3,
46
+ "rescale_factor": 0.00392156862745098
47
+ },
48
+ "image_seq_length": 280,
49
+ "processor_class": "Gemma4Processor",
50
+ "video_processor": {
51
+ "do_convert_rgb": true,
52
+ "do_normalize": true,
53
+ "do_rescale": true,
54
+ "do_resize": true,
55
+ "do_sample_frames": true,
56
+ "image_mean": [
57
+ 0.0,
58
+ 0.0,
59
+ 0.0
60
+ ],
61
+ "image_std": [
62
+ 1.0,
63
+ 1.0,
64
+ 1.0
65
+ ],
66
+ "max_soft_tokens": 70,
67
+ "num_frames": 32,
68
+ "patch_size": 16,
69
+ "pooling_kernel_size": 3,
70
+ "resample": 3,
71
+ "rescale_factor": 0.00392156862745098,
72
+ "return_metadata": false,
73
+ "video_processor_type": "Gemma4VideoProcessor"
74
+ }
75
+ }
reports/adaptive_validation_summary.json ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "settings": {
3
+ "threshold": 0.8,
4
+ "mix": 0.5,
5
+ "think_budget": 8192
6
+ },
7
+ "datasets": {
8
+ "AQuA-RAT": {
9
+ "n": 254,
10
+ "system1": 68.11023622047244,
11
+ "adaptive": 88.9763779527559,
12
+ "always_think": 90.15748031496062,
13
+ "escalated_pct": 53.937007874015755,
14
+ "ece_system1": 0.03542993923235089,
15
+ "ece_adaptive": 0.10937879489051755
16
+ },
17
+ "CommonsenseQA": {
18
+ "n": 300,
19
+ "system1": 86.66666666666667,
20
+ "adaptive": 85.0,
21
+ "always_think": 83.33333333333334,
22
+ "escalated_pct": 32.0,
23
+ "ece_system1": 0.04021547925836765,
24
+ "ece_adaptive": 0.053498085540242254
25
+ },
26
+ "LogiQA": {
27
+ "n": 300,
28
+ "system1": 55.333333333333336,
29
+ "adaptive": 80.33333333333333,
30
+ "always_think": 82.0,
31
+ "escalated_pct": 63.66666666666667,
32
+ "ece_system1": 0.12319165112240366,
33
+ "ece_adaptive": 0.12393589949576163
34
+ },
35
+ "MedMCQA": {
36
+ "n": 300,
37
+ "system1": 66.66666666666666,
38
+ "adaptive": 71.66666666666667,
39
+ "always_think": 73.66666666666667,
40
+ "escalated_pct": 52.33333333333333,
41
+ "ece_system1": 0.09009478448570694,
42
+ "ece_adaptive": 0.11489545268113895
43
+ },
44
+ "OpenBookQA": {
45
+ "n": 300,
46
+ "system1": 94.33333333333334,
47
+ "adaptive": 96.33333333333334,
48
+ "always_think": 95.66666666666667,
49
+ "escalated_pct": 14.666666666666666,
50
+ "ece_system1": 0.0443799860205645,
51
+ "ece_adaptive": 0.024977894380833043
52
+ },
53
+ "StrategyQA": {
54
+ "n": 300,
55
+ "system1": 68.0,
56
+ "adaptive": 78.66666666666666,
57
+ "always_think": 79.0,
58
+ "escalated_pct": 71.0,
59
+ "ece_system1": 0.02861885470538586,
60
+ "ece_adaptive": 0.0359552437712111
61
+ },
62
+ "ALL": {
63
+ "n": 1754,
64
+ "system1": 73.31812998859749,
65
+ "adaptive": 83.35233751425314,
66
+ "always_think": 83.80843785632838,
67
+ "escalated_pct": 47.776510832383124,
68
+ "ece_system1": 0.03522315989905164,
69
+ "ece_adaptive": 0.03458824415825564
70
+ },
71
+ "thinking_tokens": {
72
+ "median": 2282.0,
73
+ "p90": 7573.4,
74
+ "finished_within_8192": 0.9156214367160775
75
+ }
76
+ }
77
+ }
reports/decision_index_benchmark_summary.json ADDED
@@ -0,0 +1,1445 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "engine": "jev-gemma4-step7000",
3
+ "counts": {
4
+ "ok": 150759
5
+ },
6
+ "successful_request_latency_ms": {
7
+ "median": 18.467556714313105,
8
+ "p95": 215.19578801235184,
9
+ "mean": 72.46806872574511
10
+ },
11
+ "benchmarks": [
12
+ {
13
+ "catalog_id": 1,
14
+ "dataset": "BFCL",
15
+ "requests": 1694,
16
+ "answered": 1694,
17
+ "unsupported": 0,
18
+ "errors": 0,
19
+ "abstained": 0,
20
+ "pending": 0,
21
+ "scored_requests": 1694,
22
+ "metric": "case exact accuracy",
23
+ "score": 0.9462809917355371,
24
+ "reference_same_cases": null,
25
+ "median_ms": 138.82611655571964,
26
+ "detail": {
27
+ "field_accuracy": 0.9790268456375839,
28
+ "case_exact_accuracy": 0.9462809917355371,
29
+ "custom_metrics": {},
30
+ "positive_micro_f1": 0.9711649365628604
31
+ }
32
+ },
33
+ {
34
+ "catalog_id": 2,
35
+ "dataset": "ToolRet",
36
+ "requests": 685,
37
+ "answered": 685,
38
+ "unsupported": 0,
39
+ "errors": 0,
40
+ "abstained": 0,
41
+ "pending": 0,
42
+ "scored_requests": 685,
43
+ "metric": "nDCG@10",
44
+ "score": 0.6707486501882342,
45
+ "reference_same_cases": null,
46
+ "median_ms": 1291.542490303982,
47
+ "detail": {
48
+ "field_accuracy": null,
49
+ "case_exact_accuracy": null,
50
+ "custom_metrics": {
51
+ "ndcg_at_10": 0.6707486501882342,
52
+ "mrr": 0.7067817063550459,
53
+ "recall_at_10": 0.7906100104275287,
54
+ "candidate_recall": 0.8401355578727842,
55
+ "scorable_candidate_recall": 0.8401355578727842,
56
+ "candidates_scored": 31.9985401459854,
57
+ "candidates_retrieved": 32,
58
+ "bm25_ndcg_at_10": 0.5430661225020075
59
+ }
60
+ }
61
+ },
62
+ {
63
+ "catalog_id": 3,
64
+ "dataset": "API-Bank",
65
+ "requests": 508,
66
+ "answered": 508,
67
+ "unsupported": 0,
68
+ "errors": 0,
69
+ "abstained": 0,
70
+ "pending": 0,
71
+ "scored_requests": 508,
72
+ "metric": "accuracy",
73
+ "score": 0.84251968503937,
74
+ "reference_same_cases": null,
75
+ "median_ms": 1822.737945985864,
76
+ "detail": {
77
+ "field_accuracy": 0.84251968503937,
78
+ "case_exact_accuracy": 0.84251968503937,
79
+ "custom_metrics": {}
80
+ }
81
+ },
82
+ {
83
+ "catalog_id": 4,
84
+ "dataset": "BANKING77",
85
+ "requests": 3080,
86
+ "answered": 3080,
87
+ "unsupported": 0,
88
+ "errors": 0,
89
+ "abstained": 0,
90
+ "pending": 0,
91
+ "scored_requests": 3080,
92
+ "metric": "macro-F1",
93
+ "score": 0.8799267700969118,
94
+ "reference_same_cases": null,
95
+ "median_ms": 47.48895212833304,
96
+ "detail": {
97
+ "field_accuracy": 0.8821428571428571,
98
+ "case_exact_accuracy": 0.8821428571428571,
99
+ "macro_f1": 0.8799267700969118,
100
+ "custom_metrics": {}
101
+ }
102
+ },
103
+ {
104
+ "catalog_id": 5,
105
+ "dataset": "CLINC150+OOS",
106
+ "requests": 5500,
107
+ "answered": 5500,
108
+ "unsupported": 0,
109
+ "errors": 0,
110
+ "abstained": 0,
111
+ "pending": 0,
112
+ "scored_requests": 5500,
113
+ "metric": "macro-F1",
114
+ "score": 0.9297695107658135,
115
+ "reference_same_cases": null,
116
+ "median_ms": 66.9773818081012,
117
+ "detail": {
118
+ "field_accuracy": 0.9089090909090909,
119
+ "case_exact_accuracy": 0.9089090909090909,
120
+ "macro_f1": 0.9297695107658135,
121
+ "custom_metrics": {}
122
+ }
123
+ },
124
+ {
125
+ "catalog_id": 6,
126
+ "dataset": "RouterBench",
127
+ "requests": 10000,
128
+ "answered": 10000,
129
+ "unsupported": 0,
130
+ "errors": 0,
131
+ "abstained": 0,
132
+ "pending": 0,
133
+ "scored_requests": 10000,
134
+ "metric": "selected quality (quality objective)",
135
+ "score": 0.799375,
136
+ "reference_same_cases": null,
137
+ "median_ms": 168.35474733670708,
138
+ "tracks": {
139
+ "RouterBench-0shot": {
140
+ "metric": "selected quality (quality objective)",
141
+ "score": 0.7865280831501099,
142
+ "scored_requests": 5003
143
+ },
144
+ "RouterBench-5shot": {
145
+ "metric": "selected quality (quality objective)",
146
+ "score": 0.8122373424054433,
147
+ "scored_requests": 4997
148
+ }
149
+ },
150
+ "detail": {
151
+ "field_accuracy": null,
152
+ "case_exact_accuracy": null,
153
+ "custom_metrics": {
154
+ "quality_quality": 0.799375,
155
+ "quality_cost_usd": 0.00553126286,
156
+ "quality_utility": 0.799375,
157
+ "quality_oracle_optimal": 0.8312,
158
+ "quality_utility_regret": 0.119685,
159
+ "cost_aware_quality": 0.779235,
160
+ "cost_aware_cost_usd": 0.0039796064,
161
+ "cost_aware_utility": 0.739438936,
162
+ "cost_aware_oracle_optimal": 0.0314,
163
+ "cost_aware_utility_regret": 0.176090406628
164
+ }
165
+ }
166
+ },
167
+ {
168
+ "catalog_id": 9,
169
+ "dataset": "Home appliance simulator",
170
+ "requests": 88,
171
+ "answered": 88,
172
+ "unsupported": 0,
173
+ "errors": 0,
174
+ "abstained": 0,
175
+ "pending": 0,
176
+ "scored_requests": 88,
177
+ "metric": "case exact accuracy",
178
+ "score": 0.2159090909090909,
179
+ "reference_same_cases": null,
180
+ "median_ms": 569.5601852930849,
181
+ "detail": {
182
+ "field_accuracy": 0.9345911949685535,
183
+ "case_exact_accuracy": 0.2159090909090909,
184
+ "custom_metrics": {},
185
+ "subgroups": {
186
+ "outcome": {
187
+ "n": 88,
188
+ "mean": 0.8068181818181818
189
+ },
190
+ "post_0": {
191
+ "n": 69,
192
+ "mean": 0.855072463768116
193
+ },
194
+ "post_1": {
195
+ "n": 25,
196
+ "mean": 0.84
197
+ },
198
+ "resolution": {
199
+ "n": 88,
200
+ "mean": 0.7727272727272727
201
+ },
202
+ "state_read": {
203
+ "n": 88,
204
+ "mean": 1
205
+ },
206
+ "target_member": {
207
+ "n": 1232,
208
+ "mean": 0.9569805194805194
209
+ }
210
+ }
211
+ }
212
+ },
213
+ {
214
+ "catalog_id": 10,
215
+ "dataset": "SGD/SGD-X",
216
+ "requests": 2500,
217
+ "answered": 2500,
218
+ "unsupported": 0,
219
+ "errors": 0,
220
+ "abstained": 0,
221
+ "pending": 0,
222
+ "scored_requests": 2500,
223
+ "metric": "macro-F1",
224
+ "score": 0.32021537489440627,
225
+ "reference_same_cases": null,
226
+ "median_ms": 48.37857997335959,
227
+ "detail": {
228
+ "field_accuracy": 0.2868,
229
+ "case_exact_accuracy": 0.2868,
230
+ "macro_f1": 0.32021537489440627,
231
+ "custom_metrics": {}
232
+ }
233
+ },
234
+ {
235
+ "catalog_id": 11,
236
+ "dataset": "ContractNLI",
237
+ "requests": 123,
238
+ "answered": 123,
239
+ "unsupported": 0,
240
+ "errors": 0,
241
+ "abstained": 0,
242
+ "pending": 0,
243
+ "scored_requests": 123,
244
+ "metric": "macro-F1",
245
+ "score": 0.8306451780650824,
246
+ "reference_same_cases": null,
247
+ "median_ms": 1297.6135823118966,
248
+ "detail": {
249
+ "field_accuracy": 0.8689622190339551,
250
+ "case_exact_accuracy": 0.06504065040650407,
251
+ "macro_f1": 0.8306451780650824,
252
+ "custom_metrics": {}
253
+ }
254
+ },
255
+ {
256
+ "catalog_id": 12,
257
+ "dataset": "ANLI",
258
+ "requests": 3200,
259
+ "answered": 3200,
260
+ "unsupported": 0,
261
+ "errors": 0,
262
+ "abstained": 0,
263
+ "pending": 0,
264
+ "scored_requests": 3200,
265
+ "metric": "macro-F1",
266
+ "score": 0.7033787027244137,
267
+ "reference_same_cases": null,
268
+ "median_ms": 5.992660371703096,
269
+ "detail": {
270
+ "field_accuracy": 0.7021875,
271
+ "case_exact_accuracy": 0.7021875,
272
+ "macro_f1": 0.7033787027244137,
273
+ "custom_metrics": {}
274
+ }
275
+ },
276
+ {
277
+ "catalog_id": 20,
278
+ "dataset": "BPoMP",
279
+ "requests": 5000,
280
+ "answered": 5000,
281
+ "unsupported": 0,
282
+ "errors": 0,
283
+ "abstained": 0,
284
+ "pending": 0,
285
+ "scored_requests": 5000,
286
+ "metric": "accuracy",
287
+ "score": 0.9026,
288
+ "reference_same_cases": null,
289
+ "median_ms": 5.784460125141777,
290
+ "detail": {
291
+ "field_accuracy": 0.9026,
292
+ "case_exact_accuracy": 0.6399506781750924,
293
+ "custom_metrics": {},
294
+ "subgroups": {
295
+ "beam search": {
296
+ "n": 401,
297
+ "mean": 0.7880299251870324
298
+ },
299
+ "corrupted rhyme": {
300
+ "n": 376,
301
+ "mean": 0.9601063829787234
302
+ },
303
+ "delayed beam search": {
304
+ "n": 137,
305
+ "mean": 0.9416058394160584
306
+ },
307
+ "delete any 1 nonrhyming word": {
308
+ "n": 376,
309
+ "mean": 0.8297872340425532
310
+ },
311
+ "delete any 1 rhyming EOL word": {
312
+ "n": 376,
313
+ "mean": 0.976063829787234
314
+ },
315
+ "delete any 1 word": {
316
+ "n": 376,
317
+ "mean": 0.8617021276595744
318
+ },
319
+ "delete any 2 nonrhyming words": {
320
+ "n": 376,
321
+ "mean": 0.8670212765957447
322
+ },
323
+ "delete any 2 rhyming EOL words": {
324
+ "n": 376,
325
+ "mean": 0.9707446808510638
326
+ },
327
+ "delete any 2 words": {
328
+ "n": 376,
329
+ "mean": 0.9202127659574468
330
+ },
331
+ "delete any 3 nonrhyming words": {
332
+ "n": 376,
333
+ "mean": 0.8882978723404256
334
+ },
335
+ "delete any 3 rhyming words": {
336
+ "n": 376,
337
+ "mean": 0.9867021276595744
338
+ },
339
+ "delete any 3 words": {
340
+ "n": 376,
341
+ "mean": 0.9414893617021277
342
+ },
343
+ "shuffled rhyming lines": {
344
+ "n": 336,
345
+ "mean": 0.8154761904761905
346
+ },
347
+ "shuffled rhyming words": {
348
+ "n": 366,
349
+ "mean": 0.912568306010929
350
+ }
351
+ }
352
+ }
353
+ },
354
+ {
355
+ "catalog_id": 21,
356
+ "dataset": "Humicroedit",
357
+ "requests": 2628,
358
+ "answered": 2628,
359
+ "unsupported": 0,
360
+ "errors": 0,
361
+ "abstained": 0,
362
+ "pending": 0,
363
+ "scored_requests": 2628,
364
+ "metric": "accuracy",
365
+ "score": 0.630517503805175,
366
+ "reference_same_cases": null,
367
+ "median_ms": 2.9752281407127157,
368
+ "detail": {
369
+ "field_accuracy": 0.630517503805175,
370
+ "case_exact_accuracy": 0.630517503805175,
371
+ "custom_metrics": {}
372
+ }
373
+ },
374
+ {
375
+ "catalog_id": 22,
376
+ "dataset": "POP909-CL",
377
+ "requests": 2000,
378
+ "answered": 2000,
379
+ "unsupported": 0,
380
+ "errors": 0,
381
+ "abstained": 0,
382
+ "pending": 0,
383
+ "scored_requests": 2000,
384
+ "metric": "accuracy",
385
+ "score": 0.1965,
386
+ "reference_same_cases": null,
387
+ "median_ms": 1043.9998766378267,
388
+ "detail": {
389
+ "field_accuracy": 0.1965,
390
+ "case_exact_accuracy": 0.1965,
391
+ "cluster_macro_accuracy": 0.1902056277056277,
392
+ "custom_metrics": {},
393
+ "subgroups": {
394
+ "A#7": {
395
+ "n": 1,
396
+ "mean": 0
397
+ },
398
+ "A#M": {
399
+ "n": 88,
400
+ "mean": 0.20454545454545456
401
+ },
402
+ "A#m": {
403
+ "n": 43,
404
+ "mean": 0.023255813953488372
405
+ },
406
+ "A#m7": {
407
+ "n": 7,
408
+ "mean": 0.42857142857142855
409
+ },
410
+ "A#maj7": {
411
+ "n": 3,
412
+ "mean": 0.6666666666666666
413
+ },
414
+ "A7": {
415
+ "n": 4,
416
+ "mean": 0
417
+ },
418
+ "AM": {
419
+ "n": 73,
420
+ "mean": 0.1095890410958904
421
+ },
422
+ "Am": {
423
+ "n": 69,
424
+ "mean": 0.5652173913043478
425
+ },
426
+ "Am7": {
427
+ "n": 8,
428
+ "mean": 0.125
429
+ },
430
+ "Amaj7": {
431
+ "n": 3,
432
+ "mean": 0.3333333333333333
433
+ },
434
+ "B7": {
435
+ "n": 2,
436
+ "mean": 0
437
+ },
438
+ "BM": {
439
+ "n": 88,
440
+ "mean": 0.045454545454545456
441
+ },
442
+ "Bdim": {
443
+ "n": 1,
444
+ "mean": 0
445
+ },
446
+ "Bhalf-dim7": {
447
+ "n": 1,
448
+ "mean": 0
449
+ },
450
+ "Bm": {
451
+ "n": 65,
452
+ "mean": 0.06153846153846154
453
+ },
454
+ "Bm7": {
455
+ "n": 6,
456
+ "mean": 0.6666666666666666
457
+ },
458
+ "Bmaj7": {
459
+ "n": 3,
460
+ "mean": 0.3333333333333333
461
+ },
462
+ "C#7": {
463
+ "n": 1,
464
+ "mean": 1
465
+ },
466
+ "C#M": {
467
+ "n": 78,
468
+ "mean": 0.21794871794871795
469
+ },
470
+ "C#aug / Faug / Aaug": {
471
+ "n": 1,
472
+ "mean": 0
473
+ },
474
+ "C#dim": {
475
+ "n": 1,
476
+ "mean": 0
477
+ },
478
+ "C#m": {
479
+ "n": 51,
480
+ "mean": 0.3333333333333333
481
+ },
482
+ "C#m7": {
483
+ "n": 8,
484
+ "mean": 0.875
485
+ },
486
+ "C#maj7": {
487
+ "n": 6,
488
+ "mean": 0.16666666666666666
489
+ },
490
+ "C#sus2 / G#sus4": {
491
+ "n": 4,
492
+ "mean": 0
493
+ },
494
+ "C#sus4 / F#sus2": {
495
+ "n": 6,
496
+ "mean": 0.5
497
+ },
498
+ "C7": {
499
+ "n": 2,
500
+ "mean": 0
501
+ },
502
+ "CM": {
503
+ "n": 99,
504
+ "mean": 0.36363636363636365
505
+ },
506
+ "Cdim": {
507
+ "n": 1,
508
+ "mean": 0
509
+ },
510
+ "Cm": {
511
+ "n": 67,
512
+ "mean": 0.26865671641791045
513
+ },
514
+ "Cm7": {
515
+ "n": 4,
516
+ "mean": 0.5
517
+ },
518
+ "Cmaj7": {
519
+ "n": 1,
520
+ "mean": 0
521
+ },
522
+ "Csus2 / Gsus4": {
523
+ "n": 9,
524
+ "mean": 0.1111111111111111
525
+ },
526
+ "Csus4 / Fsus2": {
527
+ "n": 6,
528
+ "mean": 0.16666666666666666
529
+ },
530
+ "D#7": {
531
+ "n": 7,
532
+ "mean": 0
533
+ },
534
+ "D#M": {
535
+ "n": 86,
536
+ "mean": 0.26744186046511625
537
+ },
538
+ "D#aug / Gaug / Baug": {
539
+ "n": 1,
540
+ "mean": 0
541
+ },
542
+ "D#dim": {
543
+ "n": 1,
544
+ "mean": 0
545
+ },
546
+ "D#m": {
547
+ "n": 43,
548
+ "mean": 0.20930232558139536
549
+ },
550
+ "D#m7": {
551
+ "n": 3,
552
+ "mean": 0.3333333333333333
553
+ },
554
+ "D#maj7": {
555
+ "n": 6,
556
+ "mean": 0
557
+ },
558
+ "D#sus2 / A#sus4": {
559
+ "n": 7,
560
+ "mean": 0
561
+ },
562
+ "D#sus4 / G#sus2": {
563
+ "n": 4,
564
+ "mean": 0.25
565
+ },
566
+ "D7": {
567
+ "n": 5,
568
+ "mean": 0
569
+ },
570
+ "DM": {
571
+ "n": 77,
572
+ "mean": 0.4155844155844156
573
+ },
574
+ "Dm": {
575
+ "n": 64,
576
+ "mean": 0.1875
577
+ },
578
+ "Dm7": {
579
+ "n": 8,
580
+ "mean": 0.375
581
+ },
582
+ "DmMaj7": {
583
+ "n": 1,
584
+ "mean": 1
585
+ },
586
+ "Dmaj7": {
587
+ "n": 3,
588
+ "mean": 0
589
+ },
590
+ "Dsus2 / Asus4": {
591
+ "n": 12,
592
+ "mean": 0.08333333333333333
593
+ },
594
+ "Dsus4 / Gsus2": {
595
+ "n": 5,
596
+ "mean": 0.2
597
+ },
598
+ "E7": {
599
+ "n": 2,
600
+ "mean": 0
601
+ },
602
+ "EM": {
603
+ "n": 80,
604
+ "mean": 0.3
605
+ },
606
+ "Edim": {
607
+ "n": 1,
608
+ "mean": 0
609
+ },
610
+ "Em": {
611
+ "n": 74,
612
+ "mean": 0.0945945945945946
613
+ },
614
+ "Em7": {
615
+ "n": 6,
616
+ "mean": 0.3333333333333333
617
+ },
618
+ "Emaj7": {
619
+ "n": 1,
620
+ "mean": 0
621
+ },
622
+ "Esus2 / Bsus4": {
623
+ "n": 6,
624
+ "mean": 0.16666666666666666
625
+ },
626
+ "Esus4 / Asus2": {
627
+ "n": 8,
628
+ "mean": 0.125
629
+ },
630
+ "F#7": {
631
+ "n": 2,
632
+ "mean": 0
633
+ },
634
+ "F#M": {
635
+ "n": 62,
636
+ "mean": 0.0967741935483871
637
+ },
638
+ "F#dim": {
639
+ "n": 1,
640
+ "mean": 0
641
+ },
642
+ "F#half-dim7": {
643
+ "n": 1,
644
+ "mean": 0
645
+ },
646
+ "F#m": {
647
+ "n": 58,
648
+ "mean": 0.1724137931034483
649
+ },
650
+ "F#m7": {
651
+ "n": 9,
652
+ "mean": 0.4444444444444444
653
+ },
654
+ "F#mMaj7": {
655
+ "n": 1,
656
+ "mean": 0
657
+ },
658
+ "F#maj7": {
659
+ "n": 4,
660
+ "mean": 0
661
+ },
662
+ "F#sus4 / Bsus2": {
663
+ "n": 5,
664
+ "mean": 0
665
+ },
666
+ "F7": {
667
+ "n": 2,
668
+ "mean": 0
669
+ },
670
+ "FM": {
671
+ "n": 93,
672
+ "mean": 0.22580645161290322
673
+ },
674
+ "Fm": {
675
+ "n": 55,
676
+ "mean": 0.16363636363636364
677
+ },
678
+ "Fm7": {
679
+ "n": 7,
680
+ "mean": 0.42857142857142855
681
+ },
682
+ "Fmaj7": {
683
+ "n": 4,
684
+ "mean": 0.25
685
+ },
686
+ "Fsus4 / A#sus2": {
687
+ "n": 9,
688
+ "mean": 0.1111111111111111
689
+ },
690
+ "G#7": {
691
+ "n": 6,
692
+ "mean": 0.16666666666666666
693
+ },
694
+ "G#M": {
695
+ "n": 88,
696
+ "mean": 0.03409090909090909
697
+ },
698
+ "G#m": {
699
+ "n": 50,
700
+ "mean": 0.1
701
+ },
702
+ "G#m7": {
703
+ "n": 4,
704
+ "mean": 0.5
705
+ },
706
+ "G#maj7": {
707
+ "n": 6,
708
+ "mean": 0
709
+ },
710
+ "G7": {
711
+ "n": 2,
712
+ "mean": 0
713
+ },
714
+ "GM": {
715
+ "n": 102,
716
+ "mean": 0.0196078431372549
717
+ },
718
+ "Gdim": {
719
+ "n": 1,
720
+ "mean": 0
721
+ },
722
+ "Gm": {
723
+ "n": 66,
724
+ "mean": 0.030303030303030304
725
+ },
726
+ "Gm7": {
727
+ "n": 7,
728
+ "mean": 0.7142857142857143
729
+ },
730
+ "Gmaj7": {
731
+ "n": 2,
732
+ "mean": 0
733
+ },
734
+ "NoChord": {
735
+ "n": 29,
736
+ "mean": 0.3103448275862069
737
+ },
738
+ "Other": {
739
+ "n": 3,
740
+ "mean": 0
741
+ }
742
+ }
743
+ }
744
+ },
745
+ {
746
+ "catalog_id": 23,
747
+ "dataset": "cfcolor",
748
+ "requests": 5000,
749
+ "answered": 5000,
750
+ "unsupported": 0,
751
+ "errors": 0,
752
+ "abstained": 0,
753
+ "pending": 0,
754
+ "scored_requests": 5000,
755
+ "metric": "accuracy",
756
+ "score": 0.6316,
757
+ "reference_same_cases": null,
758
+ "median_ms": 14.917532666004263,
759
+ "detail": {
760
+ "field_accuracy": 0.6316,
761
+ "case_exact_accuracy": 0.6316,
762
+ "cluster_macro_accuracy": 0.6369418544918806,
763
+ "custom_metrics": {}
764
+ }
765
+ },
766
+ {
767
+ "catalog_id": 24,
768
+ "dataset": "MMLU",
769
+ "requests": 14033,
770
+ "answered": 14033,
771
+ "unsupported": 0,
772
+ "errors": 0,
773
+ "abstained": 0,
774
+ "pending": 0,
775
+ "scored_requests": 14033,
776
+ "metric": "accuracy",
777
+ "score": 0.8351742321670348,
778
+ "reference_same_cases": null,
779
+ "median_ms": 5.900340576772578,
780
+ "detail": {
781
+ "field_accuracy": 0.8351742321670348,
782
+ "case_exact_accuracy": 0.8351742321670348,
783
+ "custom_metrics": {}
784
+ }
785
+ },
786
+ {
787
+ "catalog_id": 25,
788
+ "dataset": "GPQA Diamond",
789
+ "requests": 196,
790
+ "answered": 196,
791
+ "unsupported": 0,
792
+ "errors": 0,
793
+ "abstained": 0,
794
+ "pending": 0,
795
+ "scored_requests": 196,
796
+ "metric": "accuracy",
797
+ "score": 0.4387755102040816,
798
+ "reference_same_cases": null,
799
+ "median_ms": 21.98228452471085,
800
+ "detail": {
801
+ "field_accuracy": 0.4387755102040816,
802
+ "case_exact_accuracy": 0.4387755102040816,
803
+ "custom_metrics": {}
804
+ }
805
+ },
806
+ {
807
+ "catalog_id": 26,
808
+ "dataset": "ARC-Easy",
809
+ "requests": 2376,
810
+ "answered": 2376,
811
+ "unsupported": 0,
812
+ "errors": 0,
813
+ "abstained": 0,
814
+ "pending": 0,
815
+ "scored_requests": 2376,
816
+ "metric": "accuracy",
817
+ "score": 0.9861111111111112,
818
+ "reference_same_cases": null,
819
+ "median_ms": 4.468785729841329,
820
+ "detail": {
821
+ "field_accuracy": 0.9861111111111112,
822
+ "case_exact_accuracy": 0.9861111111111112,
823
+ "custom_metrics": {}
824
+ }
825
+ },
826
+ {
827
+ "catalog_id": 27,
828
+ "dataset": "ARC-Challenge",
829
+ "requests": 1172,
830
+ "answered": 1172,
831
+ "unsupported": 0,
832
+ "errors": 0,
833
+ "abstained": 0,
834
+ "pending": 0,
835
+ "scored_requests": 1172,
836
+ "metric": "accuracy",
837
+ "score": 0.9658703071672355,
838
+ "reference_same_cases": null,
839
+ "median_ms": 5.025674821808934,
840
+ "detail": {
841
+ "field_accuracy": 0.9658703071672355,
842
+ "case_exact_accuracy": 0.9658703071672355,
843
+ "custom_metrics": {}
844
+ }
845
+ },
846
+ {
847
+ "catalog_id": 28,
848
+ "dataset": "WinoGrande",
849
+ "requests": 1267,
850
+ "answered": 1267,
851
+ "unsupported": 0,
852
+ "errors": 0,
853
+ "abstained": 0,
854
+ "pending": 0,
855
+ "scored_requests": 1267,
856
+ "metric": "accuracy",
857
+ "score": 0.9100236779794791,
858
+ "reference_same_cases": null,
859
+ "median_ms": 2.5069129187613726,
860
+ "detail": {
861
+ "field_accuracy": 0.9100236779794791,
862
+ "case_exact_accuracy": 0.9100236779794791,
863
+ "custom_metrics": {}
864
+ }
865
+ },
866
+ {
867
+ "catalog_id": 29,
868
+ "dataset": "HellaSwag",
869
+ "requests": 10042,
870
+ "answered": 10042,
871
+ "unsupported": 0,
872
+ "errors": 0,
873
+ "abstained": 0,
874
+ "pending": 0,
875
+ "scored_requests": 10042,
876
+ "metric": "accuracy",
877
+ "score": 0.9643497311292571,
878
+ "reference_same_cases": null,
879
+ "median_ms": 9.592797941877507,
880
+ "detail": {
881
+ "field_accuracy": 0.9643497311292571,
882
+ "case_exact_accuracy": 0.9643497311292571,
883
+ "custom_metrics": {}
884
+ }
885
+ },
886
+ {
887
+ "catalog_id": 30,
888
+ "dataset": "GSM8K",
889
+ "requests": 2638,
890
+ "answered": 2638,
891
+ "unsupported": 0,
892
+ "errors": 0,
893
+ "abstained": 0,
894
+ "pending": 0,
895
+ "scored_requests": 2638,
896
+ "metric": "accuracy",
897
+ "score": 0.9757391963608795,
898
+ "reference_same_cases": null,
899
+ "median_ms": 8.005504249013029,
900
+ "tracks": {
901
+ "GSM8K-10choice": {
902
+ "metric": "accuracy",
903
+ "score": 0.9924184988627748,
904
+ "scored_requests": 1319
905
+ },
906
+ "GSM8K-4choice": {
907
+ "metric": "accuracy",
908
+ "score": 0.959059893858984,
909
+ "scored_requests": 1319
910
+ }
911
+ },
912
+ "detail": {
913
+ "field_accuracy": 0.9757391963608795,
914
+ "case_exact_accuracy": 0.9529946929492039,
915
+ "custom_metrics": {}
916
+ }
917
+ },
918
+ {
919
+ "catalog_id": 31,
920
+ "dataset": "ChessBench",
921
+ "requests": 5000,
922
+ "answered": 5000,
923
+ "unsupported": 0,
924
+ "errors": 0,
925
+ "abstained": 0,
926
+ "pending": 0,
927
+ "scored_requests": 5000,
928
+ "metric": "accuracy",
929
+ "score": 0.2398,
930
+ "reference_same_cases": null,
931
+ "median_ms": 54.638300483929925,
932
+ "detail": {
933
+ "field_accuracy": 0.2398,
934
+ "case_exact_accuracy": 0.2398,
935
+ "custom_metrics": {
936
+ "value_regret": 0.1481518338751564
937
+ }
938
+ }
939
+ },
940
+ {
941
+ "catalog_id": 32,
942
+ "dataset": "MuSR",
943
+ "requests": 752,
944
+ "answered": 752,
945
+ "unsupported": 0,
946
+ "errors": 0,
947
+ "abstained": 0,
948
+ "pending": 0,
949
+ "scored_requests": 752,
950
+ "metric": "accuracy",
951
+ "score": 0.660904255319149,
952
+ "reference_same_cases": null,
953
+ "median_ms": 38.64438508753665,
954
+ "detail": {
955
+ "field_accuracy": 0.660904255319149,
956
+ "case_exact_accuracy": 0.660904255319149,
957
+ "custom_metrics": {}
958
+ }
959
+ },
960
+ {
961
+ "catalog_id": 33,
962
+ "dataset": "SATA-Bench",
963
+ "requests": 1650,
964
+ "answered": 1650,
965
+ "unsupported": 0,
966
+ "errors": 0,
967
+ "abstained": 0,
968
+ "pending": 0,
969
+ "scored_requests": 1650,
970
+ "metric": "case exact accuracy",
971
+ "score": 0.34424242424242424,
972
+ "reference_same_cases": null,
973
+ "median_ms": 111.17707990342751,
974
+ "detail": {
975
+ "field_accuracy": 0.848939872398015,
976
+ "case_exact_accuracy": 0.34424242424242424,
977
+ "custom_metrics": {},
978
+ "positive_micro_f1": 0.8003747232158065
979
+ }
980
+ },
981
+ {
982
+ "catalog_id": 34,
983
+ "dataset": "SimpleBench",
984
+ "requests": 10,
985
+ "answered": 10,
986
+ "unsupported": 0,
987
+ "errors": 0,
988
+ "abstained": 0,
989
+ "pending": 0,
990
+ "scored_requests": 10,
991
+ "metric": "accuracy",
992
+ "score": 0.3,
993
+ "reference_same_cases": null,
994
+ "median_ms": 475.1623761985684,
995
+ "detail": {
996
+ "field_accuracy": 0.3,
997
+ "case_exact_accuracy": 0.3,
998
+ "custom_metrics": {}
999
+ }
1000
+ },
1001
+ {
1002
+ "catalog_id": 36,
1003
+ "dataset": "BRIGHT",
1004
+ "requests": 220,
1005
+ "answered": 220,
1006
+ "unsupported": 0,
1007
+ "errors": 0,
1008
+ "abstained": 0,
1009
+ "pending": 0,
1010
+ "scored_requests": 220,
1011
+ "metric": "nDCG@10",
1012
+ "score": 0.4598626258964629,
1013
+ "reference_same_cases": null,
1014
+ "median_ms": 1153.4820798406145,
1015
+ "detail": {
1016
+ "field_accuracy": null,
1017
+ "case_exact_accuracy": null,
1018
+ "custom_metrics": {
1019
+ "ndcg_at_10": 0.4598626258964629,
1020
+ "mrr": 0.6211010141712584,
1021
+ "recall_at_10": 0.4926779384624571,
1022
+ "candidate_recall": 0.5445170600851817,
1023
+ "scorable_candidate_recall": 0.5445170600851817,
1024
+ "candidates_scored": 32,
1025
+ "candidates_retrieved": 32,
1026
+ "bm25_ndcg_at_10": 0.2761244257399229
1027
+ }
1028
+ }
1029
+ },
1030
+ {
1031
+ "catalog_id": 37,
1032
+ "dataset": "Amazon ESCI",
1033
+ "requests": 5000,
1034
+ "answered": 5000,
1035
+ "unsupported": 0,
1036
+ "errors": 0,
1037
+ "abstained": 0,
1038
+ "pending": 0,
1039
+ "scored_requests": 5000,
1040
+ "metric": "macro-F1",
1041
+ "score": 0.6082055876720927,
1042
+ "reference_same_cases": null,
1043
+ "median_ms": 28.519113766378723,
1044
+ "detail": {
1045
+ "field_accuracy": 0.7424,
1046
+ "case_exact_accuracy": 0.7424,
1047
+ "macro_f1": 0.6082055876720927,
1048
+ "custom_metrics": {}
1049
+ }
1050
+ },
1051
+ {
1052
+ "catalog_id": 38,
1053
+ "dataset": "ACOS",
1054
+ "requests": 1565,
1055
+ "answered": 1565,
1056
+ "unsupported": 0,
1057
+ "errors": 0,
1058
+ "abstained": 0,
1059
+ "pending": 0,
1060
+ "scored_requests": 1565,
1061
+ "metric": "per-review F1",
1062
+ "score": 0.24538704031549788,
1063
+ "reference_same_cases": null,
1064
+ "median_ms": 239.38009215635248,
1065
+ "detail": {
1066
+ "field_accuracy": 0.9714354718306767,
1067
+ "case_exact_accuracy": 0.0475,
1068
+ "custom_metrics": {},
1069
+ "positive_micro_f1": 0.1858573216520651
1070
+ }
1071
+ },
1072
+ {
1073
+ "catalog_id": 39,
1074
+ "dataset": "FinEntity",
1075
+ "requests": 979,
1076
+ "answered": 979,
1077
+ "unsupported": 0,
1078
+ "errors": 0,
1079
+ "abstained": 0,
1080
+ "pending": 0,
1081
+ "scored_requests": 979,
1082
+ "metric": "macro-F1",
1083
+ "score": 0.8893744438777665,
1084
+ "reference_same_cases": null,
1085
+ "median_ms": 12.923071830300614,
1086
+ "detail": {
1087
+ "field_accuracy": 0.8896195396899953,
1088
+ "case_exact_accuracy": 0.8161389172625128,
1089
+ "macro_f1": 0.8893744438777665,
1090
+ "custom_metrics": {}
1091
+ }
1092
+ },
1093
+ {
1094
+ "catalog_id": 40,
1095
+ "dataset": "iSarcasmEval",
1096
+ "requests": 4600,
1097
+ "answered": 4600,
1098
+ "unsupported": 0,
1099
+ "errors": 0,
1100
+ "abstained": 0,
1101
+ "pending": 0,
1102
+ "scored_requests": 4600,
1103
+ "metric": "see separate subtask tracks",
1104
+ "score": null,
1105
+ "reference_same_cases": null,
1106
+ "median_ms": 3.3754319592844695,
1107
+ "tracks": {
1108
+ "iSarcasmEval-A-Ar": {
1109
+ "metric": "see separate subtask tracks",
1110
+ "score": null,
1111
+ "scored_requests": 1400
1112
+ },
1113
+ "iSarcasmEval-A-En": {
1114
+ "metric": "see separate subtask tracks",
1115
+ "score": null,
1116
+ "scored_requests": 1400
1117
+ },
1118
+ "iSarcasmEval-B-En": {
1119
+ "metric": "see separate subtask tracks",
1120
+ "score": null,
1121
+ "scored_requests": 1400
1122
+ },
1123
+ "iSarcasmEval-C-Ar": {
1124
+ "metric": "see separate subtask tracks",
1125
+ "score": null,
1126
+ "scored_requests": 200
1127
+ },
1128
+ "iSarcasmEval-C-En": {
1129
+ "metric": "see separate subtask tracks",
1130
+ "score": null,
1131
+ "scored_requests": 200
1132
+ }
1133
+ },
1134
+ "detail": {
1135
+ "field_accuracy": 0.9000862068965517,
1136
+ "case_exact_accuracy": 0.8076086956521739,
1137
+ "macro_f1": 0.719505986243066,
1138
+ "custom_metrics": {},
1139
+ "positive_f1_by_field": {
1140
+ "sarcastic": 0.5152439024390244,
1141
+ "sarcasm": 0.5620915032679739,
1142
+ "irony": 0.1037037037037037,
1143
+ "satire": 0.11320754716981132,
1144
+ "understatement": 0.0,
1145
+ "overstatement": 0.04395604395604396,
1146
+ "rhetorical_question": 0.14012738853503184
1147
+ },
1148
+ "category_macro_f1": 0.21119001272451274
1149
+ }
1150
+ },
1151
+ {
1152
+ "catalog_id": 41,
1153
+ "dataset": "VAST",
1154
+ "requests": 3006,
1155
+ "answered": 3006,
1156
+ "unsupported": 0,
1157
+ "errors": 0,
1158
+ "abstained": 0,
1159
+ "pending": 0,
1160
+ "scored_requests": 3006,
1161
+ "metric": "macro-F1",
1162
+ "score": 0.6086087530434999,
1163
+ "reference_same_cases": null,
1164
+ "median_ms": 8.309710407047532,
1165
+ "detail": {
1166
+ "field_accuracy": 0.6087824351297405,
1167
+ "case_exact_accuracy": 0.6087824351297405,
1168
+ "macro_f1": 0.6086087530434999,
1169
+ "custom_metrics": {}
1170
+ }
1171
+ },
1172
+ {
1173
+ "catalog_id": 42,
1174
+ "dataset": "NLI4CT",
1175
+ "requests": 5500,
1176
+ "answered": 5500,
1177
+ "unsupported": 0,
1178
+ "errors": 0,
1179
+ "abstained": 0,
1180
+ "pending": 0,
1181
+ "scored_requests": 5500,
1182
+ "metric": "macro-F1",
1183
+ "score": 0.7734587842091947,
1184
+ "reference_same_cases": null,
1185
+ "median_ms": 37.59161247580778,
1186
+ "detail": {
1187
+ "field_accuracy": 0.8001818181818182,
1188
+ "case_exact_accuracy": 0.8001818181818182,
1189
+ "macro_f1": 0.7734587842091947,
1190
+ "custom_metrics": {}
1191
+ }
1192
+ },
1193
+ {
1194
+ "catalog_id": 43,
1195
+ "dataset": "CRUXEval",
1196
+ "requests": 570,
1197
+ "answered": 570,
1198
+ "unsupported": 0,
1199
+ "errors": 0,
1200
+ "abstained": 0,
1201
+ "pending": 0,
1202
+ "scored_requests": 570,
1203
+ "metric": "accuracy",
1204
+ "score": 0.6719298245614035,
1205
+ "reference_same_cases": null,
1206
+ "median_ms": 11.753883198252879,
1207
+ "detail": {
1208
+ "field_accuracy": 0.6719298245614035,
1209
+ "case_exact_accuracy": 0.6719298245614035,
1210
+ "custom_metrics": {}
1211
+ }
1212
+ },
1213
+ {
1214
+ "catalog_id": 44,
1215
+ "dataset": "CLadder",
1216
+ "requests": 5000,
1217
+ "answered": 5000,
1218
+ "unsupported": 0,
1219
+ "errors": 0,
1220
+ "abstained": 0,
1221
+ "pending": 0,
1222
+ "scored_requests": 5000,
1223
+ "metric": "accuracy",
1224
+ "score": 0.716,
1225
+ "reference_same_cases": null,
1226
+ "median_ms": 7.669424390769564,
1227
+ "detail": {
1228
+ "field_accuracy": 0.716,
1229
+ "case_exact_accuracy": 0.716,
1230
+ "custom_metrics": {}
1231
+ }
1232
+ },
1233
+ {
1234
+ "catalog_id": 45,
1235
+ "dataset": "HLE",
1236
+ "requests": 501,
1237
+ "answered": 501,
1238
+ "unsupported": 0,
1239
+ "errors": 0,
1240
+ "abstained": 0,
1241
+ "pending": 0,
1242
+ "scored_requests": 501,
1243
+ "metric": "accuracy",
1244
+ "score": 0.08582834331337326,
1245
+ "reference_same_cases": null,
1246
+ "median_ms": 38.43188349856064,
1247
+ "detail": {
1248
+ "field_accuracy": 0.08582834331337326,
1249
+ "case_exact_accuracy": 0.08582834331337326,
1250
+ "custom_metrics": {}
1251
+ }
1252
+ },
1253
+ {
1254
+ "catalog_id": 48,
1255
+ "dataset": "ForecastBench",
1256
+ "requests": 10139,
1257
+ "answered": 10139,
1258
+ "unsupported": 0,
1259
+ "errors": 0,
1260
+ "abstained": 0,
1261
+ "pending": 0,
1262
+ "scored_requests": 10139,
1263
+ "metric": "Brier (lower is better)",
1264
+ "score": 0.1974167741612724,
1265
+ "reference_same_cases": null,
1266
+ "median_ms": 29.15050310548395,
1267
+ "detail": {
1268
+ "field_accuracy": null,
1269
+ "case_exact_accuracy": null,
1270
+ "custom_metrics": {
1271
+ "brier": 0.1974167741612724,
1272
+ "log_loss": 0.5803510075214148,
1273
+ "accuracy": 0.7050004931452806
1274
+ },
1275
+ "subgroups": {
1276
+ "joint": {
1277
+ "n": 3242,
1278
+ "mean": 0.18390476990829577
1279
+ },
1280
+ "scalar": {
1281
+ "n": 6897,
1282
+ "mean": 0.2037682193966139
1283
+ }
1284
+ }
1285
+ }
1286
+ },
1287
+ {
1288
+ "catalog_id": 50,
1289
+ "dataset": "Habermas Machine",
1290
+ "requests": 1676,
1291
+ "answered": 1676,
1292
+ "unsupported": 0,
1293
+ "errors": 0,
1294
+ "abstained": 0,
1295
+ "pending": 0,
1296
+ "scored_requests": 1676,
1297
+ "metric": "accuracy",
1298
+ "score": 0.6079952267303103,
1299
+ "reference_same_cases": null,
1300
+ "median_ms": 36.67889098869637,
1301
+ "detail": {
1302
+ "field_accuracy": 0.6079952267303103,
1303
+ "case_exact_accuracy": 0.6079952267303103,
1304
+ "custom_metrics": {
1305
+ "normalized_rank_regret": 0.09493834526650756
1306
+ }
1307
+ }
1308
+ },
1309
+ {
1310
+ "catalog_id": 56,
1311
+ "dataset": "PhishNChips phishing decisions",
1312
+ "requests": 2000,
1313
+ "answered": 2000,
1314
+ "unsupported": 0,
1315
+ "errors": 0,
1316
+ "abstained": 0,
1317
+ "pending": 0,
1318
+ "metric": "accuracy",
1319
+ "score": 0.789,
1320
+ "median_ms": 71.1229310836643,
1321
+ "scored_requests": 2000,
1322
+ "detail": {
1323
+ "field_accuracy": 0.789,
1324
+ "scored_fields": 2000,
1325
+ "chance_on_rows": 0.5
1326
+ }
1327
+ },
1328
+ {
1329
+ "catalog_id": 57,
1330
+ "dataset": "MMLU-Pro",
1331
+ "requests": 12032,
1332
+ "answered": 12032,
1333
+ "unsupported": 0,
1334
+ "errors": 0,
1335
+ "abstained": 0,
1336
+ "pending": 0,
1337
+ "metric": "accuracy",
1338
+ "score": 0.651845079787234,
1339
+ "median_ms": 15.152119543927256,
1340
+ "scored_requests": 12032,
1341
+ "detail": {
1342
+ "field_accuracy": 0.651845079787234,
1343
+ "scored_fields": 12032,
1344
+ "chance_on_rows": 0.11087694718844986
1345
+ }
1346
+ },
1347
+ {
1348
+ "catalog_id": 58,
1349
+ "dataset": "BBH fixed-option tasks",
1350
+ "requests": 5507,
1351
+ "answered": 5507,
1352
+ "unsupported": 0,
1353
+ "errors": 0,
1354
+ "abstained": 0,
1355
+ "pending": 0,
1356
+ "metric": "accuracy",
1357
+ "score": 0.7517704739422553,
1358
+ "median_ms": 7.826623215805739,
1359
+ "scored_requests": 5507,
1360
+ "detail": {
1361
+ "field_accuracy": 0.7517704739422553,
1362
+ "scored_fields": 5507,
1363
+ "chance_on_rows": 0.31010949680343713
1364
+ }
1365
+ },
1366
+ {
1367
+ "catalog_id": 59,
1368
+ "dataset": "RAGTruth response-level hallucination",
1369
+ "requests": 2700,
1370
+ "answered": 2700,
1371
+ "unsupported": 0,
1372
+ "errors": 0,
1373
+ "abstained": 0,
1374
+ "pending": 0,
1375
+ "metric": "F1 on hallucinated class",
1376
+ "score": 0.8413865546218487,
1377
+ "median_ms": 37.95317618642002,
1378
+ "scored_requests": 2700,
1379
+ "detail": {
1380
+ "field_accuracy": 0.8881481481481481,
1381
+ "scored_fields": 2700,
1382
+ "chance_on_rows": 0.41125163541212384
1383
+ }
1384
+ },
1385
+ {
1386
+ "catalog_id": 61,
1387
+ "dataset": "HoVer claim verification",
1388
+ "requests": 4000,
1389
+ "answered": 4000,
1390
+ "unsupported": 0,
1391
+ "errors": 0,
1392
+ "abstained": 0,
1393
+ "pending": 0,
1394
+ "metric": "accuracy",
1395
+ "score": 0.87875,
1396
+ "median_ms": 20.43018453696277,
1397
+ "scored_requests": 4000,
1398
+ "detail": {
1399
+ "field_accuracy": 0.87875,
1400
+ "scored_fields": 4000,
1401
+ "chance_on_rows": 0.5
1402
+ }
1403
+ },
1404
+ {
1405
+ "catalog_id": 62,
1406
+ "dataset": "When2Call MCQ",
1407
+ "requests": 3652,
1408
+ "answered": 3652,
1409
+ "unsupported": 0,
1410
+ "errors": 0,
1411
+ "abstained": 0,
1412
+ "pending": 0,
1413
+ "metric": "accuracy",
1414
+ "score": 0.8567907995618839,
1415
+ "median_ms": 44.64486979122739,
1416
+ "scored_requests": 3652,
1417
+ "detail": {
1418
+ "field_accuracy": 0.8567907995618839,
1419
+ "scored_fields": 3652,
1420
+ "chance_on_rows": 0.25
1421
+ }
1422
+ },
1423
+ {
1424
+ "catalog_id": 64,
1425
+ "dataset": "New Yorker caption matching",
1426
+ "requests": 528,
1427
+ "answered": 528,
1428
+ "unsupported": 0,
1429
+ "errors": 0,
1430
+ "abstained": 0,
1431
+ "pending": 0,
1432
+ "metric": "accuracy",
1433
+ "score": 0.8200757575757576,
1434
+ "median_ms": 7.580137331387959,
1435
+ "scored_requests": 528,
1436
+ "detail": {
1437
+ "field_accuracy": 0.8200757575757576,
1438
+ "scored_fields": 528,
1439
+ "chance_on_rows": 0.2
1440
+ }
1441
+ }
1442
+ ],
1443
+ "note": "Native benchmark metrics on complete supported case groups. Unsupported cases excluded from accuracy but retained in coverage.",
1444
+ "edition": "0.2.1"
1445
+ }
reports/decision_index_index.json ADDED
@@ -0,0 +1,506 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "edition": "0.2.1",
3
+ "panel_id": "decision-index-0.2.1",
4
+ "index": 58.05,
5
+ "raw_index": 67.35,
6
+ "scores": {
7
+ "balanced_skill": 58.05,
8
+ "balanced_raw": 67.35,
9
+ "breadth_skill": 56.98
10
+ },
11
+ "areas": [
12
+ {
13
+ "id": "knowledge",
14
+ "label": "Knowledge & Reasoning",
15
+ "raw": 0.5484,
16
+ "skill": 0.4304,
17
+ "coverage": 1.0,
18
+ "n": 10,
19
+ "benchmarks": [
20
+ 25,
21
+ 30,
22
+ 31,
23
+ 32,
24
+ 33,
25
+ 43,
26
+ 44,
27
+ 45,
28
+ 57,
29
+ 58
30
+ ]
31
+ },
32
+ {
33
+ "id": "language",
34
+ "label": "Language Understanding",
35
+ "raw": 0.7439,
36
+ "skill": 0.6358,
37
+ "coverage": 1.0,
38
+ "n": 10,
39
+ "benchmarks": [
40
+ 11,
41
+ 12,
42
+ 28,
43
+ 29,
44
+ 38,
45
+ 39,
46
+ 40,
47
+ 41,
48
+ 42,
49
+ 59
50
+ ]
51
+ },
52
+ {
53
+ "id": "retrieval",
54
+ "label": "Retrieval & Classification",
55
+ "raw": 0.7575,
56
+ "skill": 0.6788,
57
+ "coverage": 1.0,
58
+ "n": 6,
59
+ "benchmarks": [
60
+ 4,
61
+ 5,
62
+ 36,
63
+ 37,
64
+ 56,
65
+ 61
66
+ ]
67
+ },
68
+ {
69
+ "id": "tools",
70
+ "label": "Tools & Automation",
71
+ "raw": 0.7204,
72
+ "skill": 0.6973,
73
+ "coverage": 1.0,
74
+ "n": 5,
75
+ "benchmarks": [
76
+ 1,
77
+ 2,
78
+ 3,
79
+ 9,
80
+ 62
81
+ ]
82
+ },
83
+ {
84
+ "id": "arts",
85
+ "label": "Arts & Human Taste",
86
+ "raw": 0.5615,
87
+ "skill": 0.4147,
88
+ "coverage": 1.0,
89
+ "n": 7,
90
+ "benchmarks": [
91
+ 20,
92
+ 21,
93
+ 22,
94
+ 23,
95
+ 48,
96
+ 50,
97
+ 64
98
+ ]
99
+ }
100
+ ],
101
+ "benchmarks": {
102
+ "1": {
103
+ "raw": 0.9463,
104
+ "skill": 0.9275,
105
+ "coverage": 1.0,
106
+ "random": 0.2592,
107
+ "rule": "track",
108
+ "in_index": true,
109
+ "tracks": []
110
+ },
111
+ "2": {
112
+ "raw": 0.6707,
113
+ "skill": 0.6198,
114
+ "coverage": 1.0,
115
+ "random": 0.1341,
116
+ "rule": "track",
117
+ "in_index": true,
118
+ "tracks": []
119
+ },
120
+ "3": {
121
+ "raw": 0.8425,
122
+ "skill": 0.8395,
123
+ "coverage": 1.0,
124
+ "random": 0.0189,
125
+ "rule": "chance",
126
+ "in_index": true
127
+ },
128
+ "4": {
129
+ "raw": 0.8799,
130
+ "skill": 0.8784,
131
+ "coverage": 1.0,
132
+ "random": 0.0127,
133
+ "rule": "chance",
134
+ "in_index": true
135
+ },
136
+ "5": {
137
+ "raw": 0.9298,
138
+ "skill": 0.9294,
139
+ "coverage": 1.0,
140
+ "random": 0.006,
141
+ "rule": "chance",
142
+ "in_index": true
143
+ },
144
+ "9": {
145
+ "raw": 0.2159,
146
+ "skill": 0.2159,
147
+ "coverage": 1.0,
148
+ "random": 0.0,
149
+ "rule": "chance",
150
+ "in_index": true
151
+ },
152
+ "11": {
153
+ "raw": 0.8306,
154
+ "skill": 0.7551,
155
+ "coverage": 1.0,
156
+ "random": 0.3085,
157
+ "rule": "track",
158
+ "in_index": true,
159
+ "tracks": []
160
+ },
161
+ "12": {
162
+ "raw": 0.7034,
163
+ "skill": 0.5557,
164
+ "coverage": 1.0,
165
+ "random": 0.3324,
166
+ "rule": "chance",
167
+ "in_index": true
168
+ },
169
+ "20": {
170
+ "raw": 0.9043,
171
+ "skill": 0.8085,
172
+ "coverage": 1.0,
173
+ "random": 0.5,
174
+ "rule": "track",
175
+ "in_index": true,
176
+ "tracks": []
177
+ },
178
+ "21": {
179
+ "raw": 0.6305,
180
+ "skill": 0.261,
181
+ "coverage": 1.0,
182
+ "random": 0.5,
183
+ "rule": "track",
184
+ "in_index": true,
185
+ "tracks": []
186
+ },
187
+ "22": {
188
+ "raw": 0.1902,
189
+ "skill": 0.1839,
190
+ "coverage": 1.0,
191
+ "random": 0.0078,
192
+ "rule": "track",
193
+ "in_index": true,
194
+ "tracks": []
195
+ },
196
+ "23": {
197
+ "raw": 0.6369,
198
+ "skill": 0.2739,
199
+ "coverage": 1.0,
200
+ "random": 0.5,
201
+ "rule": "track",
202
+ "in_index": true,
203
+ "tracks": []
204
+ },
205
+ "25": {
206
+ "raw": 0.4388,
207
+ "skill": 0.2517,
208
+ "coverage": 1.0,
209
+ "random": 0.25,
210
+ "rule": "track",
211
+ "in_index": true,
212
+ "tracks": []
213
+ },
214
+ "28": {
215
+ "raw": 0.91,
216
+ "skill": 0.82,
217
+ "coverage": 1.0,
218
+ "random": 0.5,
219
+ "rule": "chance",
220
+ "in_index": true
221
+ },
222
+ "29": {
223
+ "raw": 0.9643,
224
+ "skill": 0.9524,
225
+ "coverage": 1.0,
226
+ "random": 0.25,
227
+ "rule": "chance",
228
+ "in_index": true
229
+ },
230
+ "30": {
231
+ "raw": 0.9757,
232
+ "skill": 0.9685,
233
+ "coverage": 1.0,
234
+ "random": 0.25,
235
+ "rule": "track",
236
+ "in_index": true,
237
+ "tracks": [
238
+ {
239
+ "track": "GSM8K-4choice",
240
+ "score": 0.9591,
241
+ "headline": false
242
+ },
243
+ {
244
+ "track": "GSM8K-10choice",
245
+ "score": 0.9924,
246
+ "headline": false
247
+ }
248
+ ]
249
+ },
250
+ "31": {
251
+ "raw": 0.2398,
252
+ "skill": 0.172,
253
+ "coverage": 1.0,
254
+ "random": 0.0819,
255
+ "rule": "track",
256
+ "in_index": true,
257
+ "tracks": []
258
+ },
259
+ "32": {
260
+ "raw": 0.6609,
261
+ "skill": 0.4609,
262
+ "coverage": 1.0,
263
+ "random": 0.371,
264
+ "rule": "chance",
265
+ "in_index": true
266
+ },
267
+ "33": {
268
+ "raw": 0.3442,
269
+ "skill": 0.3355,
270
+ "coverage": 1.0,
271
+ "random": 0.0131,
272
+ "rule": "chance",
273
+ "in_index": true
274
+ },
275
+ "36": {
276
+ "raw": 0.4599,
277
+ "skill": 0.389,
278
+ "coverage": 1.0,
279
+ "random": 0.116,
280
+ "rule": "track",
281
+ "in_index": true,
282
+ "tracks": []
283
+ },
284
+ "37": {
285
+ "raw": 0.6082,
286
+ "skill": 0.5086,
287
+ "coverage": 1.0,
288
+ "random": 0.2027,
289
+ "rule": "track",
290
+ "in_index": true,
291
+ "tracks": []
292
+ },
293
+ "38": {
294
+ "raw": 0.2454,
295
+ "skill": 0.2213,
296
+ "coverage": 1.0,
297
+ "random": 0.031,
298
+ "rule": "chance",
299
+ "in_index": true
300
+ },
301
+ "39": {
302
+ "raw": 0.8894,
303
+ "skill": 0.8373,
304
+ "coverage": 1.0,
305
+ "random": 0.3201,
306
+ "rule": "chance",
307
+ "in_index": true
308
+ },
309
+ "40": {
310
+ "raw": 0.6027,
311
+ "skill": 0.4888,
312
+ "coverage": 1.0,
313
+ "random": 0.2227,
314
+ "rule": "track",
315
+ "in_index": true,
316
+ "tracks": [
317
+ {
318
+ "track": "A · Arabic",
319
+ "score": 0.3986,
320
+ "headline": false
321
+ },
322
+ {
323
+ "track": "A · English",
324
+ "score": 0.6027,
325
+ "headline": true
326
+ },
327
+ {
328
+ "track": "C · Arabic pairs",
329
+ "score": 0.74,
330
+ "headline": false
331
+ },
332
+ {
333
+ "track": "C · English pairs",
334
+ "score": 0.935,
335
+ "headline": false
336
+ }
337
+ ]
338
+ },
339
+ "41": {
340
+ "raw": 0.6086,
341
+ "skill": 0.4129,
342
+ "coverage": 1.0,
343
+ "random": 0.3333,
344
+ "rule": "track",
345
+ "in_index": true,
346
+ "tracks": []
347
+ },
348
+ "42": {
349
+ "raw": 0.7735,
350
+ "skill": 0.5593,
351
+ "coverage": 1.0,
352
+ "random": 0.486,
353
+ "rule": "chance",
354
+ "in_index": true
355
+ },
356
+ "43": {
357
+ "raw": 0.6719,
358
+ "skill": 0.4795,
359
+ "coverage": 1.0,
360
+ "random": 0.3697,
361
+ "rule": "track",
362
+ "in_index": true,
363
+ "tracks": []
364
+ },
365
+ "44": {
366
+ "raw": 0.716,
367
+ "skill": 0.432,
368
+ "coverage": 1.0,
369
+ "random": 0.5,
370
+ "rule": "track",
371
+ "in_index": true,
372
+ "tracks": []
373
+ },
374
+ "45": {
375
+ "raw": 0.0858,
376
+ "skill": 0.0,
377
+ "coverage": 1.0,
378
+ "random": 0.1641,
379
+ "rule": "chance",
380
+ "in_index": true
381
+ },
382
+ "48": {
383
+ "raw": 0.2104,
384
+ "skill": 0.2104,
385
+ "coverage": 1.0,
386
+ "random": 0.25,
387
+ "rule": "vs baseline",
388
+ "in_index": true
389
+ },
390
+ "50": {
391
+ "raw": 0.608,
392
+ "skill": 0.4311,
393
+ "coverage": 1.0,
394
+ "random": 0.311,
395
+ "rule": "track",
396
+ "in_index": true,
397
+ "tracks": []
398
+ },
399
+ "56": {
400
+ "raw": 0.789,
401
+ "skill": 0.578,
402
+ "coverage": 1.0,
403
+ "random": 0.5,
404
+ "rule": "chance",
405
+ "in_index": true
406
+ },
407
+ "57": {
408
+ "raw": 0.6518,
409
+ "skill": 0.6084,
410
+ "coverage": 1.0,
411
+ "random": 0.1109,
412
+ "rule": "chance",
413
+ "in_index": true
414
+ },
415
+ "58": {
416
+ "raw": 0.7518,
417
+ "skill": 0.6402,
418
+ "coverage": 1.0,
419
+ "random": 0.3101,
420
+ "rule": "chance",
421
+ "in_index": true
422
+ },
423
+ "59": {
424
+ "raw": 0.8414,
425
+ "skill": 0.6712,
426
+ "coverage": 1.0,
427
+ "random": 0.5177,
428
+ "rule": "chance",
429
+ "in_index": true
430
+ },
431
+ "61": {
432
+ "raw": 0.8788,
433
+ "skill": 0.7576,
434
+ "coverage": 1.0,
435
+ "random": 0.5,
436
+ "rule": "chance",
437
+ "in_index": true
438
+ },
439
+ "62": {
440
+ "raw": 0.8568,
441
+ "skill": 0.8091,
442
+ "coverage": 1.0,
443
+ "random": 0.25,
444
+ "rule": "chance",
445
+ "in_index": true
446
+ },
447
+ "64": {
448
+ "raw": 0.8201,
449
+ "skill": 0.7751,
450
+ "coverage": 1.0,
451
+ "random": 0.2,
452
+ "rule": "chance",
453
+ "in_index": true
454
+ },
455
+ "6": {
456
+ "raw": 0.7994,
457
+ "skill": 0.5297,
458
+ "coverage": 1.0,
459
+ "random": 0.5735,
460
+ "rule": "shown, not counted",
461
+ "in_index": false
462
+ },
463
+ "10": {
464
+ "raw": 0.3202,
465
+ "skill": 0.0,
466
+ "coverage": 1.0,
467
+ "random": 0.399,
468
+ "rule": "shown, not counted",
469
+ "in_index": false
470
+ },
471
+ "24": {
472
+ "raw": 0.8352,
473
+ "skill": 0.7803,
474
+ "coverage": 1.0,
475
+ "random": 0.25,
476
+ "rule": "shown, not counted",
477
+ "in_index": false
478
+ },
479
+ "26": {
480
+ "raw": 0.9861,
481
+ "skill": 0.9815,
482
+ "coverage": 1.0,
483
+ "random": 0.2502,
484
+ "rule": "shown, not counted",
485
+ "in_index": false
486
+ },
487
+ "27": {
488
+ "raw": 0.9659,
489
+ "skill": 0.9545,
490
+ "coverage": 1.0,
491
+ "random": 0.2502,
492
+ "rule": "shown, not counted",
493
+ "in_index": false
494
+ },
495
+ "34": {
496
+ "raw": 0.3,
497
+ "skill": 0.16,
498
+ "coverage": 1.0,
499
+ "random": 0.1667,
500
+ "rule": "shown, not counted",
501
+ "in_index": false
502
+ }
503
+ },
504
+ "coverage": 1.0,
505
+ "note": "Decision Index 0.2.1 averages 38 benchmarks in five areas. Arts & Human Taste weighs 10%; the other four share 90% in proportion to the square root of their benchmark count (knowledge 25.8%, language 25.8%, retrieval 20.0%, tools 18.3%). Inside an area, gold ★ benchmarks weigh 1.2 and the rest 1.0; the index is 100 x the weighted mean of the five areas. Each benchmark is chance-corrected first, (score - chance) / (1 - chance) clipped to 0-1, so 0 means random guessing and 100 means perfect. Every score is coverage-adjusted, so an unanswered or unsupported request counts as wrong. ForecastBench enters against its baseline: clip((0.25 - Brier) / 0.25) x coverage, so always predicting 0.5 scores zero. MMLU, ARC-Easy, ARC-Challenge, RouterBench, SGD stay on the board as non-index benchmarks. The six interactive environments are still unrun and stay out. Every entrant on the board has results on all 38 index benchmarks. Point estimates only, no uncertainty intervals yet."
506
+ }
reports/decision_index_score_brief.json ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "edition": "0.2.1",
3
+ "decision_index": 58.05,
4
+ "scores": {
5
+ "balanced_skill": 58.05,
6
+ "balanced_raw": 67.35,
7
+ "breadth_skill": 56.98
8
+ },
9
+ "areas": [
10
+ {
11
+ "id": "knowledge",
12
+ "skill": 0.4304,
13
+ "raw": 0.5484,
14
+ "coverage": 1.0,
15
+ "n": 10
16
+ },
17
+ {
18
+ "id": "language",
19
+ "skill": 0.6358,
20
+ "raw": 0.7439,
21
+ "coverage": 1.0,
22
+ "n": 10
23
+ },
24
+ {
25
+ "id": "retrieval",
26
+ "skill": 0.6788,
27
+ "raw": 0.7575,
28
+ "coverage": 1.0,
29
+ "n": 6
30
+ },
31
+ {
32
+ "id": "tools",
33
+ "skill": 0.6973,
34
+ "raw": 0.7204,
35
+ "coverage": 1.0,
36
+ "n": 5
37
+ },
38
+ {
39
+ "id": "arts",
40
+ "skill": 0.4147,
41
+ "raw": 0.5615,
42
+ "coverage": 1.0,
43
+ "n": 7
44
+ }
45
+ ],
46
+ "completed": 150317,
47
+ "complete": true,
48
+ "out": "/root/di-runs-fast/jev-gemma4-step7000"
49
+ }
reports/decision_index_scores.json ADDED
@@ -0,0 +1,1448 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "engine": "jev-gemma4-step7000",
3
+ "edition": "0.2.1",
4
+ "generated_utc": "2026-09-29T10:54:53+00:00",
5
+ "suite": {
6
+ "edition": "release-v2.1",
7
+ "requests": 120340,
8
+ "scoreable": 119898,
9
+ "excluded": 442,
10
+ "added_requests": 30419,
11
+ "benchmarks": 44,
12
+ "rows_sha256": "b2b56d6fb636837ca469e689087bdbf373dda8de7638aa2da6793e6eda0792d5",
13
+ "added_sha256": "7429f3c9cdddb772c1cfc42bb2a45e8516b0032152b746e6929f1c8b52f4ce89"
14
+ },
15
+ "completed": 150317,
16
+ "complete": true,
17
+ "counts": {
18
+ "ok": 150317
19
+ },
20
+ "latency_ms": {
21
+ "median": 18.5,
22
+ "p95": 215.2,
23
+ "mean": 72.5
24
+ },
25
+ "decision_index": 58.05,
26
+ "raw_index": 67.35,
27
+ "scores": {
28
+ "balanced_skill": 58.05,
29
+ "balanced_raw": 67.35,
30
+ "breadth_skill": 56.98
31
+ },
32
+ "areas": [
33
+ {
34
+ "id": "knowledge",
35
+ "label": "Knowledge & Reasoning",
36
+ "raw": 0.5484,
37
+ "skill": 0.4304,
38
+ "coverage": 1.0,
39
+ "n": 10,
40
+ "benchmarks": [
41
+ 25,
42
+ 30,
43
+ 31,
44
+ 32,
45
+ 33,
46
+ 43,
47
+ 44,
48
+ 45,
49
+ 57,
50
+ 58
51
+ ]
52
+ },
53
+ {
54
+ "id": "language",
55
+ "label": "Language Understanding",
56
+ "raw": 0.7439,
57
+ "skill": 0.6358,
58
+ "coverage": 1.0,
59
+ "n": 10,
60
+ "benchmarks": [
61
+ 11,
62
+ 12,
63
+ 28,
64
+ 29,
65
+ 38,
66
+ 39,
67
+ 40,
68
+ 41,
69
+ 42,
70
+ 59
71
+ ]
72
+ },
73
+ {
74
+ "id": "retrieval",
75
+ "label": "Retrieval & Classification",
76
+ "raw": 0.7575,
77
+ "skill": 0.6788,
78
+ "coverage": 1.0,
79
+ "n": 6,
80
+ "benchmarks": [
81
+ 4,
82
+ 5,
83
+ 36,
84
+ 37,
85
+ 56,
86
+ 61
87
+ ]
88
+ },
89
+ {
90
+ "id": "tools",
91
+ "label": "Tools & Automation",
92
+ "raw": 0.7204,
93
+ "skill": 0.6973,
94
+ "coverage": 1.0,
95
+ "n": 5,
96
+ "benchmarks": [
97
+ 1,
98
+ 2,
99
+ 3,
100
+ 9,
101
+ 62
102
+ ]
103
+ },
104
+ {
105
+ "id": "arts",
106
+ "label": "Arts & Human Taste",
107
+ "raw": 0.5615,
108
+ "skill": 0.4147,
109
+ "coverage": 1.0,
110
+ "n": 7,
111
+ "benchmarks": [
112
+ 20,
113
+ 21,
114
+ 22,
115
+ 23,
116
+ 48,
117
+ 50,
118
+ 64
119
+ ]
120
+ }
121
+ ],
122
+ "index_benchmarks": {
123
+ "1": {
124
+ "raw": 0.9463,
125
+ "skill": 0.9275,
126
+ "coverage": 1.0,
127
+ "random": 0.2592,
128
+ "rule": "track",
129
+ "in_index": true,
130
+ "tracks": []
131
+ },
132
+ "2": {
133
+ "raw": 0.6707,
134
+ "skill": 0.6198,
135
+ "coverage": 1.0,
136
+ "random": 0.1341,
137
+ "rule": "track",
138
+ "in_index": true,
139
+ "tracks": []
140
+ },
141
+ "3": {
142
+ "raw": 0.8425,
143
+ "skill": 0.8395,
144
+ "coverage": 1.0,
145
+ "random": 0.0189,
146
+ "rule": "chance",
147
+ "in_index": true
148
+ },
149
+ "4": {
150
+ "raw": 0.8799,
151
+ "skill": 0.8784,
152
+ "coverage": 1.0,
153
+ "random": 0.0127,
154
+ "rule": "chance",
155
+ "in_index": true
156
+ },
157
+ "5": {
158
+ "raw": 0.9298,
159
+ "skill": 0.9294,
160
+ "coverage": 1.0,
161
+ "random": 0.006,
162
+ "rule": "chance",
163
+ "in_index": true
164
+ },
165
+ "9": {
166
+ "raw": 0.2159,
167
+ "skill": 0.2159,
168
+ "coverage": 1.0,
169
+ "random": 0.0,
170
+ "rule": "chance",
171
+ "in_index": true
172
+ },
173
+ "11": {
174
+ "raw": 0.8306,
175
+ "skill": 0.7551,
176
+ "coverage": 1.0,
177
+ "random": 0.3085,
178
+ "rule": "track",
179
+ "in_index": true,
180
+ "tracks": []
181
+ },
182
+ "12": {
183
+ "raw": 0.7034,
184
+ "skill": 0.5557,
185
+ "coverage": 1.0,
186
+ "random": 0.3324,
187
+ "rule": "chance",
188
+ "in_index": true
189
+ },
190
+ "20": {
191
+ "raw": 0.9043,
192
+ "skill": 0.8085,
193
+ "coverage": 1.0,
194
+ "random": 0.5,
195
+ "rule": "track",
196
+ "in_index": true,
197
+ "tracks": []
198
+ },
199
+ "21": {
200
+ "raw": 0.6305,
201
+ "skill": 0.261,
202
+ "coverage": 1.0,
203
+ "random": 0.5,
204
+ "rule": "track",
205
+ "in_index": true,
206
+ "tracks": []
207
+ },
208
+ "22": {
209
+ "raw": 0.1902,
210
+ "skill": 0.1839,
211
+ "coverage": 1.0,
212
+ "random": 0.0078,
213
+ "rule": "track",
214
+ "in_index": true,
215
+ "tracks": []
216
+ },
217
+ "23": {
218
+ "raw": 0.6369,
219
+ "skill": 0.2739,
220
+ "coverage": 1.0,
221
+ "random": 0.5,
222
+ "rule": "track",
223
+ "in_index": true,
224
+ "tracks": []
225
+ },
226
+ "25": {
227
+ "raw": 0.4388,
228
+ "skill": 0.2517,
229
+ "coverage": 1.0,
230
+ "random": 0.25,
231
+ "rule": "track",
232
+ "in_index": true,
233
+ "tracks": []
234
+ },
235
+ "28": {
236
+ "raw": 0.91,
237
+ "skill": 0.82,
238
+ "coverage": 1.0,
239
+ "random": 0.5,
240
+ "rule": "chance",
241
+ "in_index": true
242
+ },
243
+ "29": {
244
+ "raw": 0.9643,
245
+ "skill": 0.9524,
246
+ "coverage": 1.0,
247
+ "random": 0.25,
248
+ "rule": "chance",
249
+ "in_index": true
250
+ },
251
+ "30": {
252
+ "raw": 0.9757,
253
+ "skill": 0.9685,
254
+ "coverage": 1.0,
255
+ "random": 0.25,
256
+ "rule": "track",
257
+ "in_index": true,
258
+ "tracks": [
259
+ {
260
+ "track": "GSM8K-4choice",
261
+ "score": 0.9591,
262
+ "headline": false
263
+ },
264
+ {
265
+ "track": "GSM8K-10choice",
266
+ "score": 0.9924,
267
+ "headline": false
268
+ }
269
+ ]
270
+ },
271
+ "31": {
272
+ "raw": 0.2398,
273
+ "skill": 0.172,
274
+ "coverage": 1.0,
275
+ "random": 0.0819,
276
+ "rule": "track",
277
+ "in_index": true,
278
+ "tracks": []
279
+ },
280
+ "32": {
281
+ "raw": 0.6609,
282
+ "skill": 0.4609,
283
+ "coverage": 1.0,
284
+ "random": 0.371,
285
+ "rule": "chance",
286
+ "in_index": true
287
+ },
288
+ "33": {
289
+ "raw": 0.3442,
290
+ "skill": 0.3355,
291
+ "coverage": 1.0,
292
+ "random": 0.0131,
293
+ "rule": "chance",
294
+ "in_index": true
295
+ },
296
+ "36": {
297
+ "raw": 0.4599,
298
+ "skill": 0.389,
299
+ "coverage": 1.0,
300
+ "random": 0.116,
301
+ "rule": "track",
302
+ "in_index": true,
303
+ "tracks": []
304
+ },
305
+ "37": {
306
+ "raw": 0.6082,
307
+ "skill": 0.5086,
308
+ "coverage": 1.0,
309
+ "random": 0.2027,
310
+ "rule": "track",
311
+ "in_index": true,
312
+ "tracks": []
313
+ },
314
+ "38": {
315
+ "raw": 0.2454,
316
+ "skill": 0.2213,
317
+ "coverage": 1.0,
318
+ "random": 0.031,
319
+ "rule": "chance",
320
+ "in_index": true
321
+ },
322
+ "39": {
323
+ "raw": 0.8894,
324
+ "skill": 0.8373,
325
+ "coverage": 1.0,
326
+ "random": 0.3201,
327
+ "rule": "chance",
328
+ "in_index": true
329
+ },
330
+ "40": {
331
+ "raw": 0.6027,
332
+ "skill": 0.4888,
333
+ "coverage": 1.0,
334
+ "random": 0.2227,
335
+ "rule": "track",
336
+ "in_index": true,
337
+ "tracks": [
338
+ {
339
+ "track": "A · Arabic",
340
+ "score": 0.3986,
341
+ "headline": false
342
+ },
343
+ {
344
+ "track": "A · English",
345
+ "score": 0.6027,
346
+ "headline": true
347
+ },
348
+ {
349
+ "track": "C · Arabic pairs",
350
+ "score": 0.74,
351
+ "headline": false
352
+ },
353
+ {
354
+ "track": "C · English pairs",
355
+ "score": 0.935,
356
+ "headline": false
357
+ }
358
+ ]
359
+ },
360
+ "41": {
361
+ "raw": 0.6086,
362
+ "skill": 0.4129,
363
+ "coverage": 1.0,
364
+ "random": 0.3333,
365
+ "rule": "track",
366
+ "in_index": true,
367
+ "tracks": []
368
+ },
369
+ "42": {
370
+ "raw": 0.7735,
371
+ "skill": 0.5593,
372
+ "coverage": 1.0,
373
+ "random": 0.486,
374
+ "rule": "chance",
375
+ "in_index": true
376
+ },
377
+ "43": {
378
+ "raw": 0.6719,
379
+ "skill": 0.4795,
380
+ "coverage": 1.0,
381
+ "random": 0.3697,
382
+ "rule": "track",
383
+ "in_index": true,
384
+ "tracks": []
385
+ },
386
+ "44": {
387
+ "raw": 0.716,
388
+ "skill": 0.432,
389
+ "coverage": 1.0,
390
+ "random": 0.5,
391
+ "rule": "track",
392
+ "in_index": true,
393
+ "tracks": []
394
+ },
395
+ "45": {
396
+ "raw": 0.0858,
397
+ "skill": 0.0,
398
+ "coverage": 1.0,
399
+ "random": 0.1641,
400
+ "rule": "chance",
401
+ "in_index": true
402
+ },
403
+ "48": {
404
+ "raw": 0.2104,
405
+ "skill": 0.2104,
406
+ "coverage": 1.0,
407
+ "random": 0.25,
408
+ "rule": "vs baseline",
409
+ "in_index": true
410
+ },
411
+ "50": {
412
+ "raw": 0.608,
413
+ "skill": 0.4311,
414
+ "coverage": 1.0,
415
+ "random": 0.311,
416
+ "rule": "track",
417
+ "in_index": true,
418
+ "tracks": []
419
+ },
420
+ "56": {
421
+ "raw": 0.789,
422
+ "skill": 0.578,
423
+ "coverage": 1.0,
424
+ "random": 0.5,
425
+ "rule": "chance",
426
+ "in_index": true
427
+ },
428
+ "57": {
429
+ "raw": 0.6518,
430
+ "skill": 0.6084,
431
+ "coverage": 1.0,
432
+ "random": 0.1109,
433
+ "rule": "chance",
434
+ "in_index": true
435
+ },
436
+ "58": {
437
+ "raw": 0.7518,
438
+ "skill": 0.6402,
439
+ "coverage": 1.0,
440
+ "random": 0.3101,
441
+ "rule": "chance",
442
+ "in_index": true
443
+ },
444
+ "59": {
445
+ "raw": 0.8414,
446
+ "skill": 0.6712,
447
+ "coverage": 1.0,
448
+ "random": 0.5177,
449
+ "rule": "chance",
450
+ "in_index": true
451
+ },
452
+ "61": {
453
+ "raw": 0.8788,
454
+ "skill": 0.7576,
455
+ "coverage": 1.0,
456
+ "random": 0.5,
457
+ "rule": "chance",
458
+ "in_index": true
459
+ },
460
+ "62": {
461
+ "raw": 0.8568,
462
+ "skill": 0.8091,
463
+ "coverage": 1.0,
464
+ "random": 0.25,
465
+ "rule": "chance",
466
+ "in_index": true
467
+ },
468
+ "64": {
469
+ "raw": 0.8201,
470
+ "skill": 0.7751,
471
+ "coverage": 1.0,
472
+ "random": 0.2,
473
+ "rule": "chance",
474
+ "in_index": true
475
+ },
476
+ "6": {
477
+ "raw": 0.7994,
478
+ "skill": 0.5297,
479
+ "coverage": 1.0,
480
+ "random": 0.5735,
481
+ "rule": "shown, not counted",
482
+ "in_index": false
483
+ },
484
+ "10": {
485
+ "raw": 0.3202,
486
+ "skill": 0.0,
487
+ "coverage": 1.0,
488
+ "random": 0.399,
489
+ "rule": "shown, not counted",
490
+ "in_index": false
491
+ },
492
+ "24": {
493
+ "raw": 0.8352,
494
+ "skill": 0.7803,
495
+ "coverage": 1.0,
496
+ "random": 0.25,
497
+ "rule": "shown, not counted",
498
+ "in_index": false
499
+ },
500
+ "26": {
501
+ "raw": 0.9861,
502
+ "skill": 0.9815,
503
+ "coverage": 1.0,
504
+ "random": 0.2502,
505
+ "rule": "shown, not counted",
506
+ "in_index": false
507
+ },
508
+ "27": {
509
+ "raw": 0.9659,
510
+ "skill": 0.9545,
511
+ "coverage": 1.0,
512
+ "random": 0.2502,
513
+ "rule": "shown, not counted",
514
+ "in_index": false
515
+ },
516
+ "34": {
517
+ "raw": 0.3,
518
+ "skill": 0.16,
519
+ "coverage": 1.0,
520
+ "random": 0.1667,
521
+ "rule": "shown, not counted",
522
+ "in_index": false
523
+ }
524
+ },
525
+ "benchmarks": {
526
+ "1": {
527
+ "catalog_id": 1,
528
+ "dataset": "BFCL",
529
+ "requests": 1694,
530
+ "answered": 1694,
531
+ "unsupported": 0,
532
+ "errors": 0,
533
+ "abstained": 0,
534
+ "pending": 0,
535
+ "scored_requests": 1694,
536
+ "metric": "case exact accuracy",
537
+ "score": 0.9463,
538
+ "reference_same_cases": null,
539
+ "median_ms": 138.8,
540
+ "index_raw": 0.9463,
541
+ "index_skill": 0.9275,
542
+ "coverage": 1.0,
543
+ "chance": 0.2592,
544
+ "in_index": true
545
+ },
546
+ "2": {
547
+ "catalog_id": 2,
548
+ "dataset": "ToolRet",
549
+ "requests": 685,
550
+ "answered": 685,
551
+ "unsupported": 0,
552
+ "errors": 0,
553
+ "abstained": 0,
554
+ "pending": 0,
555
+ "scored_requests": 685,
556
+ "metric": "nDCG@10",
557
+ "score": 0.6707,
558
+ "reference_same_cases": null,
559
+ "median_ms": 1291.5,
560
+ "index_raw": 0.6707,
561
+ "index_skill": 0.6198,
562
+ "coverage": 1.0,
563
+ "chance": 0.1341,
564
+ "in_index": true
565
+ },
566
+ "3": {
567
+ "catalog_id": 3,
568
+ "dataset": "API-Bank",
569
+ "requests": 508,
570
+ "answered": 508,
571
+ "unsupported": 0,
572
+ "errors": 0,
573
+ "abstained": 0,
574
+ "pending": 0,
575
+ "scored_requests": 508,
576
+ "metric": "accuracy",
577
+ "score": 0.8425,
578
+ "reference_same_cases": null,
579
+ "median_ms": 1822.7,
580
+ "index_raw": 0.8425,
581
+ "index_skill": 0.8395,
582
+ "coverage": 1.0,
583
+ "chance": 0.0189,
584
+ "in_index": true
585
+ },
586
+ "4": {
587
+ "catalog_id": 4,
588
+ "dataset": "BANKING77",
589
+ "requests": 3080,
590
+ "answered": 3080,
591
+ "unsupported": 0,
592
+ "errors": 0,
593
+ "abstained": 0,
594
+ "pending": 0,
595
+ "scored_requests": 3080,
596
+ "metric": "macro-F1",
597
+ "score": 0.8799,
598
+ "reference_same_cases": null,
599
+ "median_ms": 47.5,
600
+ "index_raw": 0.8799,
601
+ "index_skill": 0.8784,
602
+ "coverage": 1.0,
603
+ "chance": 0.0127,
604
+ "in_index": true
605
+ },
606
+ "5": {
607
+ "catalog_id": 5,
608
+ "dataset": "CLINC150+OOS",
609
+ "requests": 5500,
610
+ "answered": 5500,
611
+ "unsupported": 0,
612
+ "errors": 0,
613
+ "abstained": 0,
614
+ "pending": 0,
615
+ "scored_requests": 5500,
616
+ "metric": "macro-F1",
617
+ "score": 0.9298,
618
+ "reference_same_cases": null,
619
+ "median_ms": 67.0,
620
+ "index_raw": 0.9298,
621
+ "index_skill": 0.9294,
622
+ "coverage": 1.0,
623
+ "chance": 0.006,
624
+ "in_index": true
625
+ },
626
+ "6": {
627
+ "catalog_id": 6,
628
+ "dataset": "RouterBench",
629
+ "requests": 10000,
630
+ "answered": 10000,
631
+ "unsupported": 0,
632
+ "errors": 0,
633
+ "abstained": 0,
634
+ "pending": 0,
635
+ "scored_requests": 10000,
636
+ "metric": "selected quality (quality objective)",
637
+ "score": 0.7994,
638
+ "reference_same_cases": null,
639
+ "median_ms": 168.4,
640
+ "tracks": {
641
+ "RouterBench-0shot": {
642
+ "metric": "selected quality (quality objective)",
643
+ "score": 0.7865280831501099,
644
+ "scored_requests": 5003
645
+ },
646
+ "RouterBench-5shot": {
647
+ "metric": "selected quality (quality objective)",
648
+ "score": 0.8122373424054433,
649
+ "scored_requests": 4997
650
+ }
651
+ },
652
+ "index_raw": 0.7994,
653
+ "index_skill": 0.5297,
654
+ "coverage": 1.0,
655
+ "chance": 0.5735,
656
+ "in_index": false
657
+ },
658
+ "9": {
659
+ "catalog_id": 9,
660
+ "dataset": "Home appliance simulator",
661
+ "requests": 88,
662
+ "answered": 88,
663
+ "unsupported": 0,
664
+ "errors": 0,
665
+ "abstained": 0,
666
+ "pending": 0,
667
+ "scored_requests": 88,
668
+ "metric": "case exact accuracy",
669
+ "score": 0.2159,
670
+ "reference_same_cases": null,
671
+ "median_ms": 569.6,
672
+ "index_raw": 0.2159,
673
+ "index_skill": 0.2159,
674
+ "coverage": 1.0,
675
+ "chance": 0.0,
676
+ "in_index": true
677
+ },
678
+ "10": {
679
+ "catalog_id": 10,
680
+ "dataset": "SGD/SGD-X",
681
+ "requests": 2500,
682
+ "answered": 2500,
683
+ "unsupported": 0,
684
+ "errors": 0,
685
+ "abstained": 0,
686
+ "pending": 0,
687
+ "scored_requests": 2500,
688
+ "metric": "macro-F1",
689
+ "score": 0.3202,
690
+ "reference_same_cases": null,
691
+ "median_ms": 48.4,
692
+ "index_raw": 0.3202,
693
+ "index_skill": 0.0,
694
+ "coverage": 1.0,
695
+ "chance": 0.399,
696
+ "in_index": false
697
+ },
698
+ "11": {
699
+ "catalog_id": 11,
700
+ "dataset": "ContractNLI",
701
+ "requests": 123,
702
+ "answered": 123,
703
+ "unsupported": 0,
704
+ "errors": 0,
705
+ "abstained": 0,
706
+ "pending": 0,
707
+ "scored_requests": 123,
708
+ "metric": "macro-F1",
709
+ "score": 0.8306,
710
+ "reference_same_cases": null,
711
+ "median_ms": 1297.6,
712
+ "index_raw": 0.8306,
713
+ "index_skill": 0.7551,
714
+ "coverage": 1.0,
715
+ "chance": 0.3085,
716
+ "in_index": true
717
+ },
718
+ "12": {
719
+ "catalog_id": 12,
720
+ "dataset": "ANLI",
721
+ "requests": 3200,
722
+ "answered": 3200,
723
+ "unsupported": 0,
724
+ "errors": 0,
725
+ "abstained": 0,
726
+ "pending": 0,
727
+ "scored_requests": 3200,
728
+ "metric": "macro-F1",
729
+ "score": 0.7034,
730
+ "reference_same_cases": null,
731
+ "median_ms": 6.0,
732
+ "index_raw": 0.7034,
733
+ "index_skill": 0.5557,
734
+ "coverage": 1.0,
735
+ "chance": 0.3324,
736
+ "in_index": true
737
+ },
738
+ "20": {
739
+ "catalog_id": 20,
740
+ "dataset": "BPoMP",
741
+ "requests": 5000,
742
+ "answered": 5000,
743
+ "unsupported": 0,
744
+ "errors": 0,
745
+ "abstained": 0,
746
+ "pending": 0,
747
+ "scored_requests": 5000,
748
+ "metric": "accuracy",
749
+ "score": 0.9026,
750
+ "reference_same_cases": null,
751
+ "median_ms": 5.8,
752
+ "index_raw": 0.9043,
753
+ "index_skill": 0.8085,
754
+ "coverage": 1.0,
755
+ "chance": 0.5,
756
+ "in_index": true
757
+ },
758
+ "21": {
759
+ "catalog_id": 21,
760
+ "dataset": "Humicroedit",
761
+ "requests": 2628,
762
+ "answered": 2628,
763
+ "unsupported": 0,
764
+ "errors": 0,
765
+ "abstained": 0,
766
+ "pending": 0,
767
+ "scored_requests": 2628,
768
+ "metric": "accuracy",
769
+ "score": 0.6305,
770
+ "reference_same_cases": null,
771
+ "median_ms": 3.0,
772
+ "index_raw": 0.6305,
773
+ "index_skill": 0.261,
774
+ "coverage": 1.0,
775
+ "chance": 0.5,
776
+ "in_index": true
777
+ },
778
+ "22": {
779
+ "catalog_id": 22,
780
+ "dataset": "POP909-CL",
781
+ "requests": 2000,
782
+ "answered": 2000,
783
+ "unsupported": 0,
784
+ "errors": 0,
785
+ "abstained": 0,
786
+ "pending": 0,
787
+ "scored_requests": 2000,
788
+ "metric": "accuracy",
789
+ "score": 0.1965,
790
+ "reference_same_cases": null,
791
+ "median_ms": 1044.0,
792
+ "index_raw": 0.1902,
793
+ "index_skill": 0.1839,
794
+ "coverage": 1.0,
795
+ "chance": 0.0078,
796
+ "in_index": true
797
+ },
798
+ "23": {
799
+ "catalog_id": 23,
800
+ "dataset": "cfcolor",
801
+ "requests": 5000,
802
+ "answered": 5000,
803
+ "unsupported": 0,
804
+ "errors": 0,
805
+ "abstained": 0,
806
+ "pending": 0,
807
+ "scored_requests": 5000,
808
+ "metric": "accuracy",
809
+ "score": 0.6316,
810
+ "reference_same_cases": null,
811
+ "median_ms": 14.9,
812
+ "index_raw": 0.6369,
813
+ "index_skill": 0.2739,
814
+ "coverage": 1.0,
815
+ "chance": 0.5,
816
+ "in_index": true
817
+ },
818
+ "24": {
819
+ "catalog_id": 24,
820
+ "dataset": "MMLU",
821
+ "requests": 14033,
822
+ "answered": 14033,
823
+ "unsupported": 0,
824
+ "errors": 0,
825
+ "abstained": 0,
826
+ "pending": 0,
827
+ "scored_requests": 14033,
828
+ "metric": "accuracy",
829
+ "score": 0.8352,
830
+ "reference_same_cases": null,
831
+ "median_ms": 5.9,
832
+ "index_raw": 0.8352,
833
+ "index_skill": 0.7803,
834
+ "coverage": 1.0,
835
+ "chance": 0.25,
836
+ "in_index": false
837
+ },
838
+ "25": {
839
+ "catalog_id": 25,
840
+ "dataset": "GPQA Diamond",
841
+ "requests": 196,
842
+ "answered": 196,
843
+ "unsupported": 0,
844
+ "errors": 0,
845
+ "abstained": 0,
846
+ "pending": 0,
847
+ "scored_requests": 196,
848
+ "metric": "accuracy",
849
+ "score": 0.4388,
850
+ "reference_same_cases": null,
851
+ "median_ms": 22.0,
852
+ "index_raw": 0.4388,
853
+ "index_skill": 0.2517,
854
+ "coverage": 1.0,
855
+ "chance": 0.25,
856
+ "in_index": true
857
+ },
858
+ "26": {
859
+ "catalog_id": 26,
860
+ "dataset": "ARC-Easy",
861
+ "requests": 2376,
862
+ "answered": 2376,
863
+ "unsupported": 0,
864
+ "errors": 0,
865
+ "abstained": 0,
866
+ "pending": 0,
867
+ "scored_requests": 2376,
868
+ "metric": "accuracy",
869
+ "score": 0.9861,
870
+ "reference_same_cases": null,
871
+ "median_ms": 4.5,
872
+ "index_raw": 0.9861,
873
+ "index_skill": 0.9815,
874
+ "coverage": 1.0,
875
+ "chance": 0.2502,
876
+ "in_index": false
877
+ },
878
+ "27": {
879
+ "catalog_id": 27,
880
+ "dataset": "ARC-Challenge",
881
+ "requests": 1172,
882
+ "answered": 1172,
883
+ "unsupported": 0,
884
+ "errors": 0,
885
+ "abstained": 0,
886
+ "pending": 0,
887
+ "scored_requests": 1172,
888
+ "metric": "accuracy",
889
+ "score": 0.9659,
890
+ "reference_same_cases": null,
891
+ "median_ms": 5.0,
892
+ "index_raw": 0.9659,
893
+ "index_skill": 0.9545,
894
+ "coverage": 1.0,
895
+ "chance": 0.2502,
896
+ "in_index": false
897
+ },
898
+ "28": {
899
+ "catalog_id": 28,
900
+ "dataset": "WinoGrande",
901
+ "requests": 1267,
902
+ "answered": 1267,
903
+ "unsupported": 0,
904
+ "errors": 0,
905
+ "abstained": 0,
906
+ "pending": 0,
907
+ "scored_requests": 1267,
908
+ "metric": "accuracy",
909
+ "score": 0.91,
910
+ "reference_same_cases": null,
911
+ "median_ms": 2.5,
912
+ "index_raw": 0.91,
913
+ "index_skill": 0.82,
914
+ "coverage": 1.0,
915
+ "chance": 0.5,
916
+ "in_index": true
917
+ },
918
+ "29": {
919
+ "catalog_id": 29,
920
+ "dataset": "HellaSwag",
921
+ "requests": 10042,
922
+ "answered": 10042,
923
+ "unsupported": 0,
924
+ "errors": 0,
925
+ "abstained": 0,
926
+ "pending": 0,
927
+ "scored_requests": 10042,
928
+ "metric": "accuracy",
929
+ "score": 0.9643,
930
+ "reference_same_cases": null,
931
+ "median_ms": 9.6,
932
+ "index_raw": 0.9643,
933
+ "index_skill": 0.9524,
934
+ "coverage": 1.0,
935
+ "chance": 0.25,
936
+ "in_index": true
937
+ },
938
+ "30": {
939
+ "catalog_id": 30,
940
+ "dataset": "GSM8K",
941
+ "requests": 2638,
942
+ "answered": 2638,
943
+ "unsupported": 0,
944
+ "errors": 0,
945
+ "abstained": 0,
946
+ "pending": 0,
947
+ "scored_requests": 2638,
948
+ "metric": "accuracy",
949
+ "score": 0.9757,
950
+ "reference_same_cases": null,
951
+ "median_ms": 8.0,
952
+ "tracks": [
953
+ {
954
+ "track": "GSM8K-4choice",
955
+ "score": 0.9591,
956
+ "headline": false
957
+ },
958
+ {
959
+ "track": "GSM8K-10choice",
960
+ "score": 0.9924,
961
+ "headline": false
962
+ }
963
+ ],
964
+ "index_raw": 0.9757,
965
+ "index_skill": 0.9685,
966
+ "coverage": 1.0,
967
+ "chance": 0.25,
968
+ "in_index": true
969
+ },
970
+ "31": {
971
+ "catalog_id": 31,
972
+ "dataset": "ChessBench",
973
+ "requests": 5000,
974
+ "answered": 5000,
975
+ "unsupported": 0,
976
+ "errors": 0,
977
+ "abstained": 0,
978
+ "pending": 0,
979
+ "scored_requests": 5000,
980
+ "metric": "accuracy",
981
+ "score": 0.2398,
982
+ "reference_same_cases": null,
983
+ "median_ms": 54.6,
984
+ "index_raw": 0.2398,
985
+ "index_skill": 0.172,
986
+ "coverage": 1.0,
987
+ "chance": 0.0819,
988
+ "in_index": true
989
+ },
990
+ "32": {
991
+ "catalog_id": 32,
992
+ "dataset": "MuSR",
993
+ "requests": 752,
994
+ "answered": 752,
995
+ "unsupported": 0,
996
+ "errors": 0,
997
+ "abstained": 0,
998
+ "pending": 0,
999
+ "scored_requests": 752,
1000
+ "metric": "accuracy",
1001
+ "score": 0.6609,
1002
+ "reference_same_cases": null,
1003
+ "median_ms": 38.6,
1004
+ "index_raw": 0.6609,
1005
+ "index_skill": 0.4609,
1006
+ "coverage": 1.0,
1007
+ "chance": 0.371,
1008
+ "in_index": true
1009
+ },
1010
+ "33": {
1011
+ "catalog_id": 33,
1012
+ "dataset": "SATA-Bench",
1013
+ "requests": 1650,
1014
+ "answered": 1650,
1015
+ "unsupported": 0,
1016
+ "errors": 0,
1017
+ "abstained": 0,
1018
+ "pending": 0,
1019
+ "scored_requests": 1650,
1020
+ "metric": "case exact accuracy",
1021
+ "score": 0.3442,
1022
+ "reference_same_cases": null,
1023
+ "median_ms": 111.2,
1024
+ "index_raw": 0.3442,
1025
+ "index_skill": 0.3355,
1026
+ "coverage": 1.0,
1027
+ "chance": 0.0131,
1028
+ "in_index": true
1029
+ },
1030
+ "34": {
1031
+ "catalog_id": 34,
1032
+ "dataset": "SimpleBench",
1033
+ "requests": 10,
1034
+ "answered": 10,
1035
+ "unsupported": 0,
1036
+ "errors": 0,
1037
+ "abstained": 0,
1038
+ "pending": 0,
1039
+ "scored_requests": 10,
1040
+ "metric": "accuracy",
1041
+ "score": 0.3,
1042
+ "reference_same_cases": null,
1043
+ "median_ms": 475.2,
1044
+ "index_raw": 0.3,
1045
+ "index_skill": 0.16,
1046
+ "coverage": 1.0,
1047
+ "chance": 0.1667,
1048
+ "in_index": false
1049
+ },
1050
+ "36": {
1051
+ "catalog_id": 36,
1052
+ "dataset": "BRIGHT",
1053
+ "requests": 220,
1054
+ "answered": 220,
1055
+ "unsupported": 0,
1056
+ "errors": 0,
1057
+ "abstained": 0,
1058
+ "pending": 0,
1059
+ "scored_requests": 220,
1060
+ "metric": "nDCG@10",
1061
+ "score": 0.4599,
1062
+ "reference_same_cases": null,
1063
+ "median_ms": 1153.5,
1064
+ "index_raw": 0.4599,
1065
+ "index_skill": 0.389,
1066
+ "coverage": 1.0,
1067
+ "chance": 0.116,
1068
+ "in_index": true
1069
+ },
1070
+ "37": {
1071
+ "catalog_id": 37,
1072
+ "dataset": "Amazon ESCI",
1073
+ "requests": 5000,
1074
+ "answered": 5000,
1075
+ "unsupported": 0,
1076
+ "errors": 0,
1077
+ "abstained": 0,
1078
+ "pending": 0,
1079
+ "scored_requests": 5000,
1080
+ "metric": "macro-F1",
1081
+ "score": 0.6082,
1082
+ "reference_same_cases": null,
1083
+ "median_ms": 28.5,
1084
+ "index_raw": 0.6082,
1085
+ "index_skill": 0.5086,
1086
+ "coverage": 1.0,
1087
+ "chance": 0.2027,
1088
+ "in_index": true
1089
+ },
1090
+ "38": {
1091
+ "catalog_id": 38,
1092
+ "dataset": "ACOS",
1093
+ "requests": 1565,
1094
+ "answered": 1565,
1095
+ "unsupported": 0,
1096
+ "errors": 0,
1097
+ "abstained": 0,
1098
+ "pending": 0,
1099
+ "scored_requests": 1565,
1100
+ "metric": "per-review F1",
1101
+ "score": 0.2454,
1102
+ "reference_same_cases": null,
1103
+ "median_ms": 239.4,
1104
+ "index_raw": 0.2454,
1105
+ "index_skill": 0.2213,
1106
+ "coverage": 1.0,
1107
+ "chance": 0.031,
1108
+ "in_index": true
1109
+ },
1110
+ "39": {
1111
+ "catalog_id": 39,
1112
+ "dataset": "FinEntity",
1113
+ "requests": 979,
1114
+ "answered": 979,
1115
+ "unsupported": 0,
1116
+ "errors": 0,
1117
+ "abstained": 0,
1118
+ "pending": 0,
1119
+ "scored_requests": 979,
1120
+ "metric": "macro-F1",
1121
+ "score": 0.8894,
1122
+ "reference_same_cases": null,
1123
+ "median_ms": 12.9,
1124
+ "index_raw": 0.8894,
1125
+ "index_skill": 0.8373,
1126
+ "coverage": 1.0,
1127
+ "chance": 0.3201,
1128
+ "in_index": true
1129
+ },
1130
+ "40": {
1131
+ "catalog_id": 40,
1132
+ "dataset": "iSarcasmEval",
1133
+ "requests": 4600,
1134
+ "answered": 4600,
1135
+ "unsupported": 0,
1136
+ "errors": 0,
1137
+ "abstained": 0,
1138
+ "pending": 0,
1139
+ "scored_requests": 4600,
1140
+ "metric": "Sarcasm F1 · track A, English",
1141
+ "score": 0.6027,
1142
+ "reference_same_cases": null,
1143
+ "median_ms": 3.4,
1144
+ "tracks": [
1145
+ {
1146
+ "track": "A · Arabic",
1147
+ "score": 0.3986,
1148
+ "headline": false
1149
+ },
1150
+ {
1151
+ "track": "A · English",
1152
+ "score": 0.6027,
1153
+ "headline": true
1154
+ },
1155
+ {
1156
+ "track": "C · Arabic pairs",
1157
+ "score": 0.74,
1158
+ "headline": false
1159
+ },
1160
+ {
1161
+ "track": "C · English pairs",
1162
+ "score": 0.935,
1163
+ "headline": false
1164
+ }
1165
+ ],
1166
+ "index_raw": 0.6027,
1167
+ "index_skill": 0.4888,
1168
+ "coverage": 1.0,
1169
+ "chance": 0.2227,
1170
+ "in_index": true
1171
+ },
1172
+ "41": {
1173
+ "catalog_id": 41,
1174
+ "dataset": "VAST",
1175
+ "requests": 3006,
1176
+ "answered": 3006,
1177
+ "unsupported": 0,
1178
+ "errors": 0,
1179
+ "abstained": 0,
1180
+ "pending": 0,
1181
+ "scored_requests": 3006,
1182
+ "metric": "macro-F1",
1183
+ "score": 0.6086,
1184
+ "reference_same_cases": null,
1185
+ "median_ms": 8.3,
1186
+ "index_raw": 0.6086,
1187
+ "index_skill": 0.4129,
1188
+ "coverage": 1.0,
1189
+ "chance": 0.3333,
1190
+ "in_index": true
1191
+ },
1192
+ "42": {
1193
+ "catalog_id": 42,
1194
+ "dataset": "NLI4CT",
1195
+ "requests": 5500,
1196
+ "answered": 5500,
1197
+ "unsupported": 0,
1198
+ "errors": 0,
1199
+ "abstained": 0,
1200
+ "pending": 0,
1201
+ "scored_requests": 5500,
1202
+ "metric": "macro-F1",
1203
+ "score": 0.7735,
1204
+ "reference_same_cases": null,
1205
+ "median_ms": 37.6,
1206
+ "index_raw": 0.7735,
1207
+ "index_skill": 0.5593,
1208
+ "coverage": 1.0,
1209
+ "chance": 0.486,
1210
+ "in_index": true
1211
+ },
1212
+ "43": {
1213
+ "catalog_id": 43,
1214
+ "dataset": "CRUXEval",
1215
+ "requests": 570,
1216
+ "answered": 570,
1217
+ "unsupported": 0,
1218
+ "errors": 0,
1219
+ "abstained": 0,
1220
+ "pending": 0,
1221
+ "scored_requests": 570,
1222
+ "metric": "accuracy",
1223
+ "score": 0.6719,
1224
+ "reference_same_cases": null,
1225
+ "median_ms": 11.8,
1226
+ "index_raw": 0.6719,
1227
+ "index_skill": 0.4795,
1228
+ "coverage": 1.0,
1229
+ "chance": 0.3697,
1230
+ "in_index": true
1231
+ },
1232
+ "44": {
1233
+ "catalog_id": 44,
1234
+ "dataset": "CLadder",
1235
+ "requests": 5000,
1236
+ "answered": 5000,
1237
+ "unsupported": 0,
1238
+ "errors": 0,
1239
+ "abstained": 0,
1240
+ "pending": 0,
1241
+ "scored_requests": 5000,
1242
+ "metric": "accuracy",
1243
+ "score": 0.716,
1244
+ "reference_same_cases": null,
1245
+ "median_ms": 7.7,
1246
+ "index_raw": 0.716,
1247
+ "index_skill": 0.432,
1248
+ "coverage": 1.0,
1249
+ "chance": 0.5,
1250
+ "in_index": true
1251
+ },
1252
+ "45": {
1253
+ "catalog_id": 45,
1254
+ "dataset": "HLE",
1255
+ "requests": 501,
1256
+ "answered": 501,
1257
+ "unsupported": 0,
1258
+ "errors": 0,
1259
+ "abstained": 0,
1260
+ "pending": 0,
1261
+ "scored_requests": 501,
1262
+ "metric": "accuracy",
1263
+ "score": 0.0858,
1264
+ "reference_same_cases": null,
1265
+ "median_ms": 38.4,
1266
+ "index_raw": 0.0858,
1267
+ "index_skill": 0.0,
1268
+ "coverage": 1.0,
1269
+ "chance": 0.1641,
1270
+ "in_index": true
1271
+ },
1272
+ "48": {
1273
+ "catalog_id": 48,
1274
+ "dataset": "ForecastBench",
1275
+ "requests": 10139,
1276
+ "answered": 10139,
1277
+ "unsupported": 0,
1278
+ "errors": 0,
1279
+ "abstained": 0,
1280
+ "pending": 0,
1281
+ "scored_requests": 10139,
1282
+ "metric": "Brier (lower is better)",
1283
+ "score": 0.1974,
1284
+ "reference_same_cases": null,
1285
+ "median_ms": 29.2,
1286
+ "index_raw": 0.2104,
1287
+ "index_skill": 0.2104,
1288
+ "coverage": 1.0,
1289
+ "chance": 0.25,
1290
+ "in_index": true
1291
+ },
1292
+ "50": {
1293
+ "catalog_id": 50,
1294
+ "dataset": "Habermas Machine",
1295
+ "requests": 1676,
1296
+ "answered": 1676,
1297
+ "unsupported": 0,
1298
+ "errors": 0,
1299
+ "abstained": 0,
1300
+ "pending": 0,
1301
+ "scored_requests": 1676,
1302
+ "metric": "accuracy",
1303
+ "score": 0.608,
1304
+ "reference_same_cases": null,
1305
+ "median_ms": 36.7,
1306
+ "index_raw": 0.608,
1307
+ "index_skill": 0.4311,
1308
+ "coverage": 1.0,
1309
+ "chance": 0.311,
1310
+ "in_index": true
1311
+ },
1312
+ "56": {
1313
+ "catalog_id": 56,
1314
+ "dataset": "PhishNChips phishing decisions",
1315
+ "requests": 2000,
1316
+ "answered": 2000,
1317
+ "unsupported": 0,
1318
+ "errors": 0,
1319
+ "abstained": 0,
1320
+ "pending": 0,
1321
+ "metric": "accuracy",
1322
+ "score": 0.789,
1323
+ "median_ms": 71.1,
1324
+ "scored_requests": 2000,
1325
+ "index_raw": 0.789,
1326
+ "index_skill": 0.578,
1327
+ "coverage": 1.0,
1328
+ "chance": 0.5,
1329
+ "in_index": true
1330
+ },
1331
+ "57": {
1332
+ "catalog_id": 57,
1333
+ "dataset": "MMLU-Pro",
1334
+ "requests": 12032,
1335
+ "answered": 12032,
1336
+ "unsupported": 0,
1337
+ "errors": 0,
1338
+ "abstained": 0,
1339
+ "pending": 0,
1340
+ "metric": "accuracy",
1341
+ "score": 0.6518,
1342
+ "median_ms": 15.2,
1343
+ "scored_requests": 12032,
1344
+ "index_raw": 0.6518,
1345
+ "index_skill": 0.6084,
1346
+ "coverage": 1.0,
1347
+ "chance": 0.1109,
1348
+ "in_index": true
1349
+ },
1350
+ "58": {
1351
+ "catalog_id": 58,
1352
+ "dataset": "BBH fixed-option tasks",
1353
+ "requests": 5507,
1354
+ "answered": 5507,
1355
+ "unsupported": 0,
1356
+ "errors": 0,
1357
+ "abstained": 0,
1358
+ "pending": 0,
1359
+ "metric": "accuracy",
1360
+ "score": 0.7518,
1361
+ "median_ms": 7.8,
1362
+ "scored_requests": 5507,
1363
+ "index_raw": 0.7518,
1364
+ "index_skill": 0.6402,
1365
+ "coverage": 1.0,
1366
+ "chance": 0.3101,
1367
+ "in_index": true
1368
+ },
1369
+ "59": {
1370
+ "catalog_id": 59,
1371
+ "dataset": "RAGTruth response-level hallucination",
1372
+ "requests": 2700,
1373
+ "answered": 2700,
1374
+ "unsupported": 0,
1375
+ "errors": 0,
1376
+ "abstained": 0,
1377
+ "pending": 0,
1378
+ "metric": "F1 on hallucinated class",
1379
+ "score": 0.8414,
1380
+ "median_ms": 38.0,
1381
+ "scored_requests": 2700,
1382
+ "index_raw": 0.8414,
1383
+ "index_skill": 0.6712,
1384
+ "coverage": 1.0,
1385
+ "chance": 0.5177,
1386
+ "in_index": true
1387
+ },
1388
+ "61": {
1389
+ "catalog_id": 61,
1390
+ "dataset": "HoVer claim verification",
1391
+ "requests": 4000,
1392
+ "answered": 4000,
1393
+ "unsupported": 0,
1394
+ "errors": 0,
1395
+ "abstained": 0,
1396
+ "pending": 0,
1397
+ "metric": "accuracy",
1398
+ "score": 0.8788,
1399
+ "median_ms": 20.4,
1400
+ "scored_requests": 4000,
1401
+ "index_raw": 0.8788,
1402
+ "index_skill": 0.7576,
1403
+ "coverage": 1.0,
1404
+ "chance": 0.5,
1405
+ "in_index": true
1406
+ },
1407
+ "62": {
1408
+ "catalog_id": 62,
1409
+ "dataset": "When2Call MCQ",
1410
+ "requests": 3652,
1411
+ "answered": 3652,
1412
+ "unsupported": 0,
1413
+ "errors": 0,
1414
+ "abstained": 0,
1415
+ "pending": 0,
1416
+ "metric": "accuracy",
1417
+ "score": 0.8568,
1418
+ "median_ms": 44.6,
1419
+ "scored_requests": 3652,
1420
+ "index_raw": 0.8568,
1421
+ "index_skill": 0.8091,
1422
+ "coverage": 1.0,
1423
+ "chance": 0.25,
1424
+ "in_index": true
1425
+ },
1426
+ "64": {
1427
+ "catalog_id": 64,
1428
+ "dataset": "New Yorker caption matching",
1429
+ "requests": 528,
1430
+ "answered": 528,
1431
+ "unsupported": 0,
1432
+ "errors": 0,
1433
+ "abstained": 0,
1434
+ "pending": 0,
1435
+ "metric": "accuracy",
1436
+ "score": 0.8201,
1437
+ "median_ms": 7.6,
1438
+ "scored_requests": 528,
1439
+ "index_raw": 0.8201,
1440
+ "index_skill": 0.7751,
1441
+ "coverage": 1.0,
1442
+ "chance": 0.2,
1443
+ "in_index": true
1444
+ }
1445
+ },
1446
+ "panel_id": "decision-index-0.2.1",
1447
+ "note": "Decision Index 0.2.1 averages 38 benchmarks in five areas. Arts & Human Taste weighs 10%; the other four share 90% in proportion to the square root of their benchmark count (knowledge 25.8%, language 25.8%, retrieval 20.0%, tools 18.3%). Inside an area, gold ★ benchmarks weigh 1.2 and the rest 1.0; the index is 100 x the weighted mean of the five areas. Each benchmark is chance-corrected first, (score - chance) / (1 - chance) clipped to 0-1, so 0 means random guessing and 100 means perfect. Every score is coverage-adjusted, so an unanswered or unsupported request counts as wrong. ForecastBench enters against its baseline: clip((0.25 - Brier) / 0.25) x coverage, so always predicting 0.5 scores zero. MMLU, ARC-Easy, ARC-Challenge, RouterBench, SGD stay on the board as non-index benchmarks. The six interactive environments are still unrun and stay out. Every entrant on the board has results on all 38 index benchmarks. Point estimates only, no uncertainty intervals yet."
1448
+ }
serve.sh ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ # Serve autotrust/GEV-26B-Decide with vLLM: one engine, both systems, text and images.
3
+ # System 2: the unmodified gemma-4-26B-A4B-it (served name autotrust/GEV-26B-Decide) on the OpenAI endpoints
4
+ # System 1: POST /v1/decide (2-256 options, optional adaptive thinking), or the LoRA module "jev-decision" directly
5
+ # serve_decide.py is the standard vLLM OpenAI server (same flags) with /v1/decide added.
6
+ # Requirements: a vLLM build with Gemma-4 support plus patches/vllm-gemma4-lm-head-lora.patch (LoRA on Gemma-4's tied
7
+ # lm_head, vocabulary 262,144), tested with a vLLM development build from September 2026.
8
+ # MAX_MODEL_LEN: up to 262144 (the backbone's native context).
9
+ set -e
10
+ MODEL_DIR=${MODEL_DIR:-GEV-26B-Decide}
11
+ [ -d "$MODEL_DIR" ] || hf download autotrust/GEV-26B-Decide --local-dir "$MODEL_DIR"
12
+ exec python3 "$MODEL_DIR/serve_decide.py" --model "$MODEL_DIR" --served-model-name autotrust/GEV-26B-Decide \
13
+ --enable-lora --max-lora-rank 32 --lora-modules jev-decision="$MODEL_DIR/adapter_vllm" \
14
+ --logprobs-mode processed_logprobs --max-model-len ${MAX_MODEL_LEN:-65536} --enable-prefix-caching \
15
+ --max-num-seqs 256 --trust-request-chat-template --limit-mm-per-prompt '{"image": 8}' \
16
+ --port ${PORT:-8000}
serve_decide.py ADDED
@@ -0,0 +1,367 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """vLLM OpenAI-compatible server with a built-in System 1 endpoint (POST /v1/decide) and adaptive thinking.
3
+
4
+ One engine serves both systems, for JEV-27B / JEV-27B-VL (Qwen3.8) and GEV-26B-A4B (Gemma-4):
5
+ System 2 the unmodified base model, through the usual OpenAI endpoints (served model name)
6
+ System 1 the LoRA module "jev-decision" (backbone LoRA + decision head as an lm_head LoRA), through /v1/decide
7
+
8
+ POST /v1/decide
9
+ kind "noul" (yes/no), "score" (0-5) or "choice" (2-256 options)
10
+ state string, JSON object, or a list mixing text and images: ["text", {"image": "https://... | data:..."}]
11
+ question string
12
+ options list of strings (choice only)
13
+ strategy choices with more than 16 options: "single" (one pass, labels A-P then Q-Z, AA, ...), "tournament"
14
+ (groups of <=16 + a final of 16), "permute" (single pass over 4 option orders, averaged); default per model
15
+ thinking "off" (default), "auto" (think only when the leading option is below `threshold`), "on" (always think)
16
+ threshold System 1 confidence below which "auto" switches thinking on (default per model)
17
+ think_budget maximum thinking tokens (default 8192)
18
+ return_reasoning include System 2's reasoning text
19
+ debug include the System 1 and System 2 distributions
20
+
21
+ Adaptive thinking: System 1 gives p1 in one pass. If thinking is switched on, the base model reasons in its thinking
22
+ mode over the same state, question and options; when the thinking channel closes, the answer-letter distribution p2 is
23
+ read in one step, and the result is p = (1 - w) * p1 + w * p2 (default w = 0.5: the fast decision and the reasoning get
24
+ equal weight, which keeps the calibration of System 1; p2 alone is close to one-hot and over-confident).
25
+
26
+ Response (same shape as the hosted JEV API, plus "thinking" when requested):
27
+ {"kind", "effective_kind", "options", "probabilities", "choice_index", "choice", "adaptation", "protocol", "model",
28
+ "usage", "elapsed_seconds", "num_model_requests", "thinking": {...}}
29
+
30
+ Run it like `python -m vllm.entrypoints.openai.api_server` (same flags), with --lora-modules jev-decision=<adapter dir>
31
+ and --trust-request-chat-template. decision_head.json is read from the adapter directory, calibration.json from its
32
+ parent (override with JEV_DECIDE_CALIBRATION).
33
+ """
34
+ import asyncio
35
+ import json
36
+ import math
37
+ import os
38
+ import string
39
+ import time
40
+ import zlib
41
+ from typing import Literal
42
+
43
+ from fastapi import APIRouter, Request
44
+ from fastapi.responses import JSONResponse
45
+ from pydantic import BaseModel
46
+
47
+ import vllm.entrypoints.launchers.api_server.entry as entry
48
+ from vllm.entrypoints.openai.chat_completion.protocol import ChatCompletionRequest
49
+ from vllm.entrypoints.openai.completion.protocol import CompletionRequest
50
+
51
+ LORA = os.environ.get("JEV_DECIDE_LORA", "jev-decision")
52
+ MAX_OPTIONS = 256
53
+ CHUNK = 128 # vLLM caps logprob_token_ids at 128 per request
54
+ PROTOCOL = "jev27-bare-v1"
55
+ PLACEHOLDER = "XQXCONTENTXQX"
56
+
57
+ # Per-model settings. threshold / mix: adaptive-thinking defaults. GEV: fitted on 1,754 questions outside the Decision
58
+ # Index (CommonsenseQA, OpenBookQA, AQuA-RAT, MedMCQA, LogiQA, StrategyQA). JEV-27B: same defaults, not yet validated.
59
+ PROFILES = {
60
+ "qwen": {"prefix": "", "image": "<|vision_start|><|image_pad|><|vision_end|>", "end_think": "</think>",
61
+ "after_think": "\n\n", "strategy": "single", "threshold": 0.8, "mix": 0.5},
62
+ "gemma": {"prefix": "<bos>", "image": "<|image|>", "end_think": "<channel|>", "after_think": "",
63
+ "strategy": "tournament", "threshold": 0.8, "mix": 0.5},
64
+ }
65
+ # Read the full distribution: override generation_config defaults (top_k/top_p) that would truncate processed logprobs.
66
+ READ = dict(max_tokens=1, temperature=1.0, top_p=1.0, top_k=0, min_p=0.0, repetition_penalty=1.0,
67
+ add_special_tokens=False, return_tokens_as_token_ids=True)
68
+ S = {}
69
+
70
+
71
+ class DecideRequest(BaseModel):
72
+ kind: Literal["noul", "score", "choice"]
73
+ state: str | dict | list = ""
74
+ question: str
75
+ options: list[str] | None = None
76
+ strategy: Literal["auto", "single", "tournament", "permute"] = "auto"
77
+ thinking: Literal["off", "auto", "on"] = "off"
78
+ threshold: float | None = None
79
+ think_budget: int = 8192
80
+ return_reasoning: bool = False
81
+ debug: bool = False
82
+
83
+
84
+ # ------------------------------------------------------------------ setup
85
+ def setup(args):
86
+ from transformers import AutoConfig, AutoTokenizer
87
+
88
+ path = next((m.path for m in (args.lora_modules or []) if m.name == LORA), None)
89
+ if path is None:
90
+ raise SystemExit(f"serve_decide: start vLLM with --lora-modules {LORA}=<adapter dir>")
91
+ head = json.load(open(os.path.join(path, "decision_head.json")))
92
+ calib = os.environ.get("JEV_DECIDE_CALIBRATION") or os.path.join(os.path.dirname(os.path.abspath(path)), "calibration.json")
93
+ temps = json.load(open(calib))["per_kind"] if os.path.exists(calib) else {"noul": 1.0, "score": 1.0, "choice": 1.0}
94
+ temps.setdefault("score", 1.0)
95
+ src = args.tokenizer or args.model
96
+ tok = AutoTokenizer.from_pretrained(src, trust_remote_code=args.trust_remote_code)
97
+ mtype = AutoConfig.from_pretrained(args.model, trust_remote_code=args.trust_remote_code).model_type
98
+ prof = dict(PROFILES["gemma" if "gemma" in mtype else "qwen"])
99
+ for k in ("strategy", "threshold", "mix"):
100
+ if os.environ.get(f"JEV_DECIDE_{k.upper()}"):
101
+ prof[k] = os.environ[f"JEV_DECIDE_{k.upper()}"] if k == "strategy" else float(os.environ[f"JEV_DECIDE_{k.upper()}"])
102
+
103
+ def single_token_labels(context):
104
+ out = []
105
+ for lab in list(string.ascii_uppercase) + [a + b for a in string.ascii_uppercase for b in string.ascii_uppercase]:
106
+ t = tok.encode(lab, add_special_tokens=False)
107
+ if len(t) == 1 and t[0] in tok.encode(context.format(lab), add_special_tokens=False):
108
+ out.append((lab, t[0]))
109
+ if len(out) == MAX_OPTIONS:
110
+ break
111
+ return out
112
+
113
+ dec = single_token_labels("x\n{}) y") # System 1 option-line labels
114
+ lo, hi = head["slots"]["ranges"]["choice"]
115
+ assert [t for _, t in dec[: hi - lo]] == head["verbalizer_ids"][lo:hi], "first labels must be the trained A-P head"
116
+ base = tok.encode("Answer: (", add_special_tokens=False) # System 2 answer labels
117
+ ans = []
118
+ for lab in list(string.ascii_uppercase) + [a + b for a in string.ascii_uppercase for b in string.ascii_uppercase]:
119
+ ids = tok.encode(f"Answer: ({lab})", add_special_tokens=False)
120
+ if ids[: len(base)] == base and len(ids) == len(base) + 2:
121
+ ans.append((lab, ids[len(base)]))
122
+ if len(ans) == MAX_OPTIONS:
123
+ break
124
+ chat = tok.apply_chat_template([{"role": "user", "content": PLACEHOLDER}], tokenize=False, add_generation_prompt=True,
125
+ enable_thinking=True)
126
+ pre, post = chat.split(PLACEHOLDER)
127
+ raw = ("{%- for m in messages -%}{%- if m['content'] is string -%}{{ m['content'] }}{%- else -%}"
128
+ "{%- for c in m['content'] -%}{%- if c['type'] == 'text' -%}{{ c['text'] }}{%- else -%}" + prof["image"] +
129
+ "{%- endif -%}{%- endfor -%}{%- endif -%}{%- endfor -%}")
130
+ names = args.served_model_name
131
+ S.update(head=head, temps=temps, labels=[l for l, _ in dec], label_ids=[t for _, t in dec], ans_labels=ans, prof=prof,
132
+ think_pre=pre, think_post=post, raw=raw, end_think_id=tok.convert_tokens_to_ids(prof["end_think"]),
133
+ model=(names[0] if isinstance(names, list) else names) or args.model, model_type=mtype)
134
+
135
+
136
+ # ------------------------------------------------------------------ helpers
137
+ def _parts(state):
138
+ if isinstance(state, str):
139
+ return [{"type": "text", "text": state}], False
140
+ if isinstance(state, dict):
141
+ return [{"type": "text", "text": json.dumps(state, ensure_ascii=False)}], False
142
+ out, img = [], False
143
+ for p in state:
144
+ if isinstance(p, str):
145
+ out.append({"type": "text", "text": p})
146
+ elif isinstance(p, dict) and "image" in p:
147
+ out.append({"type": "image_url", "image_url": {"url": p["image"]}}); img = True
148
+ elif isinstance(p, dict) and p.get("type") == "image_url":
149
+ out.append(p); img = True
150
+ elif isinstance(p, dict) and p.get("type") == "text":
151
+ out.append(p)
152
+ else:
153
+ out.append({"type": "text", "text": json.dumps(p, ensure_ascii=False)})
154
+ return out, img
155
+
156
+
157
+ def _softmax(z):
158
+ m = max(z); e = [math.exp(x - m) for x in z]; s = sum(e)
159
+ return [x / s for x in e]
160
+
161
+
162
+ def _err(msg, code=400):
163
+ return JSONResponse({"error": {"message": msg, "type": "BadRequestError", "code": code}}, status_code=code)
164
+
165
+
166
+ class Upstream(Exception):
167
+ def __init__(self, resp):
168
+ self.resp = resp
169
+
170
+
171
+ def _check(out):
172
+ if hasattr(out, "error"):
173
+ raise Upstream(out)
174
+ return out
175
+
176
+
177
+ class Ctx:
178
+ def __init__(self, raw):
179
+ self.raw, self.requests, self.prompt_tokens, self.completion_tokens = raw, 0, 0, 0
180
+
181
+ def count(self, out):
182
+ self.requests += 1
183
+ if getattr(out, "usage", None):
184
+ self.prompt_tokens += out.usage.prompt_tokens or 0
185
+ self.completion_tokens += out.usage.completion_tokens or 0
186
+
187
+
188
+ async def _logprobs(ctx, content, has_img, ids, model, allowed=None):
189
+ """One-token read-out: {token_id: logprob} for `ids` after the content (text or chat parts)."""
190
+ lp = {}
191
+ for i in range(0, len(ids), CHUNK):
192
+ chunk = ids[i: i + CHUNK]
193
+ kw = dict(logprob_token_ids=chunk) if allowed is None else dict(allowed_token_ids=allowed)
194
+ if has_img:
195
+ r = ChatCompletionRequest(model=model, messages=[{"role": "user", "content": content}], chat_template=S["raw"],
196
+ add_generation_prompt=False, logprobs=True, top_logprobs=len(chunk) if allowed else 1, **kw, **READ)
197
+ out = _check(await ctx.raw.app.state.openai_serving_chat.create_chat_completion(r, ctx.raw)); ctx.count(out)
198
+ lp.update({int(t.token.split(":")[1]): t.logprob for t in out.choices[0].logprobs.content[0].top_logprobs})
199
+ else:
200
+ text = "".join(c["text"] for c in content)
201
+ r = CompletionRequest(model=model, prompt=text, logprobs=len(chunk) if allowed else 1, **kw, **READ)
202
+ out = _check(await ctx.raw.app.state.openai_serving_completion.create_completion(r, ctx.raw)); ctx.count(out)
203
+ lp.update({int(k.split(":")[1]): v for k, v in out.choices[0].logprobs.top_logprobs[0].items()})
204
+ if allowed is not None:
205
+ break
206
+ return lp
207
+
208
+
209
+ # ------------------------------------------------------------------ System 1
210
+ async def s1_pass(ctx, kind, parts, has_img, question, opts):
211
+ """One System 1 pass over `opts` (<=256). Returns probabilities aligned with opts."""
212
+ head, temps, prof = S["head"], S["temps"], S["prof"]
213
+ lo, hi = head["slots"]["ranges"][kind]
214
+ if kind == "choice":
215
+ n = len(opts); ids = S["label_ids"][:n]
216
+ bias = [head["bias"][lo + i] if lo + i < hi else 0.0 for i in range(n)]
217
+ lines = [f"{S['labels'][i]}) {o}" for i, o in enumerate(opts)]
218
+ else:
219
+ ids = head["verbalizer_ids"][lo:hi]; bias = head["bias"][lo:hi]; lines = opts
220
+ content = ([{"type": "text", "text": f"{prof['prefix']}[kind] {kind}\n[state] "}] + parts +
221
+ [{"type": "text", "text": f"\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"}])
222
+ lp = await _logprobs(ctx, content, has_img, ids, LORA, allowed=ids if len(ids) <= 16 else None)
223
+ return _softmax([(max(lp.get(t, -1e9), -1e9) + b) / temps[kind] for t, b in zip(ids, bias)])
224
+
225
+
226
+ def _groups(n, k=16):
227
+ g = math.ceil(n / k); base, extra = divmod(n, g); out, i = [], 0
228
+ for j in range(g):
229
+ size = base + (1 if j < extra else 0); out.append(list(range(i, i + size))); i += size
230
+ return out
231
+
232
+
233
+ async def s1_dist(ctx, kind, parts, has_img, question, opts, strategy):
234
+ if kind != "choice" or len(opts) <= 16 or strategy == "single":
235
+ return await s1_pass(ctx, kind, parts, has_img, question, opts)
236
+ n = len(opts)
237
+ if strategy == "permute":
238
+ import random
239
+ rng = random.Random(zlib.crc32(question.encode()))
240
+ orders = [list(range(n))] + [rng.sample(range(n), n) for _ in range(3)]
241
+ res = await asyncio.gather(*(s1_pass(ctx, kind, parts, has_img, question, [opts[i] for i in o]) for o in orders))
242
+ p = [0.0] * n
243
+ for o, r in zip(orders, res):
244
+ for i, v in zip(o, r):
245
+ p[i] += v / len(orders)
246
+ return p
247
+ groups = _groups(n) # tournament: groups of <=16 in the given order (in parallel), then a final of 16
248
+ parts_g = await asyncio.gather(*(s1_pass(ctx, kind, parts, has_img, question, [opts[i] for i in g]) for g in groups))
249
+ in_group = {o: p for g, ps in zip(groups, parts_g) for o, p in zip(g, ps)}
250
+ chosen = [max(g, key=lambda o: (in_group[o], -o)) for g in groups]
251
+ rest = sorted((o for g in groups for o in g if o not in set(chosen)), key=lambda o: (-in_group[o], o))
252
+ fin = sorted(chosen + rest[: max(0, 16 - len(chosen))])
253
+ final = dict(zip(fin, await s1_pass(ctx, kind, parts, has_img, question, [opts[i] for i in fin])))
254
+ group_of = {o: gi for gi, g in enumerate(groups) for o in g}
255
+ share = [0.0] * len(groups); cap = [0.0] * len(groups)
256
+ for f in fin:
257
+ share[group_of[f]] += final[f]; cap[group_of[f]] += in_group[f]
258
+ among = sum(a * b for a, b in zip(share, cap))
259
+ p = [final[o] * among if o in final else share[group_of[o]] * in_group[o] for o in range(n)]
260
+ s = sum(p)
261
+ return [x / s for x in p]
262
+
263
+
264
+ # ------------------------------------------------------------------ System 2 (thinking)
265
+ async def s2_dist(ctx, kind, parts, has_img, question, opts, budget, want_text):
266
+ """The base model thinks over the same input; the answer-letter distribution is read after the thinking channel."""
267
+ shown = ["Yes (true)", "No (false)"] if kind == "noul" else opts
268
+ labs = S["ans_labels"][: len(shown)]
269
+ body = "\n".join(f"({l}) {o}" for (l, _), o in zip(labs, shown))
270
+ tail = (f"\n\nQuestion: {question}\n\nOptions:\n{body}\n\n"
271
+ "Think it through carefully, then give your final answer on the last line in the form: Answer: (X)")
272
+ user = [{"type": "text", "text": S["think_pre"]}] + parts + [{"type": "text", "text": tail + S["think_post"]}]
273
+ seed = zlib.crc32((question + body).encode()) & 0x7FFFFFFF
274
+ t0 = time.time()
275
+ gen = ChatCompletionRequest(model=S["model"], messages=[{"role": "user", "content": user}], chat_template=S["raw"],
276
+ add_generation_prompt=False, add_special_tokens=False, max_tokens=budget, seed=seed,
277
+ stop_token_ids=[S["end_think_id"]], skip_special_tokens=False)
278
+ out = _check(await ctx.raw.app.state.openai_serving_chat.create_chat_completion(gen, ctx.raw)); ctx.count(out)
279
+ thought = out.choices[0].message.content or ""
280
+ finished = out.choices[0].finish_reason == "stop"
281
+ ntok = out.usage.completion_tokens if out.usage else None
282
+ read = user + [{"type": "text", "text": thought + S["prof"]["end_think"] + S["prof"]["after_think"] + "Answer: ("}]
283
+ ids = [t for _, t in labs]
284
+ lp = await _logprobs(ctx, read, True, ids, S["model"], allowed=ids if len(ids) <= 16 else None)
285
+ p = _softmax([max(lp.get(t, -1e9), -1e9) for t in ids])
286
+ if kind == "noul":
287
+ p = [p[1], p[0]] # (A) yes / (B) no -> [P(false), P(true)]
288
+ info = {"think_tokens": ntok, "think_seconds": round(time.time() - t0, 3), "finished_within_budget": finished}
289
+ if want_text:
290
+ info["reasoning"] = thought.replace("<|channel>thought\n", "").strip()
291
+ return p, info
292
+
293
+
294
+ def _fold(p1, p2, w):
295
+ return [(1 - w) * x + w * y for x, y in zip(p1, p2)]
296
+
297
+
298
+ # ------------------------------------------------------------------ routes
299
+ router = APIRouter()
300
+
301
+
302
+ @router.get("/v1/decide/info")
303
+ async def info():
304
+ p = S["prof"]
305
+ return {"protocol": PROTOCOL, "model": S["model"], "model_type": S["model_type"], "max_options": len(S["labels"]),
306
+ "native_choice_options": len(S["head"]["slots"]["verbalizers"]) - 8, "temperatures": S["temps"],
307
+ "defaults": {"strategy": p["strategy"], "thinking": "off", "threshold": p["threshold"],
308
+ "mix": p["mix"], "think_budget": 8192}}
309
+
310
+
311
+ @router.post("/v1/decide")
312
+ async def decide(req: DecideRequest, raw: Request):
313
+ t0 = time.time()
314
+ if req.kind == "choice":
315
+ opts = req.options or []
316
+ if not 2 <= len(opts) <= len(S["labels"]):
317
+ return _err(f"choice needs 2-{len(S['labels'])} options, got {len(opts)}")
318
+ else:
319
+ opts = ["false", "true"] if req.kind == "noul" else [str(i) for i in range(6)]
320
+ if req.thinking != "off" and req.kind == "score":
321
+ return _err("thinking is supported for noul and choice")
322
+ prof = S["prof"]
323
+ strategy = prof["strategy"] if req.strategy == "auto" else req.strategy
324
+ tau = prof["threshold"] if req.threshold is None else req.threshold
325
+ parts, has_img = _parts(req.state)
326
+ ctx = Ctx(raw)
327
+ try:
328
+ p1 = await s1_dist(ctx, req.kind, parts, has_img, req.question, opts, strategy)
329
+ probs, think = p1, None
330
+ if req.thinking == "on" or (req.thinking == "auto" and max(p1) < tau):
331
+ p2, think = await s2_dist(ctx, req.kind, parts, has_img, req.question, opts, req.think_budget, req.return_reasoning)
332
+ probs = _fold(p1, p2, prof["mix"])
333
+ think = {"used": True, **think}
334
+ if req.debug:
335
+ think.update(system1=p1, system2=p2)
336
+ elif req.thinking != "off":
337
+ think = {"used": False}
338
+ if req.debug:
339
+ think["system1"] = p1
340
+ except Upstream as e:
341
+ return JSONResponse(e.resp.model_dump(), status_code=e.resp.error.code)
342
+ k = max(range(len(probs)), key=probs.__getitem__)
343
+ native = req.kind != "choice" or len(opts) <= 16
344
+ resp = {"kind": req.kind, "effective_kind": req.kind, "options": opts, "probabilities": probs, "choice_index": k,
345
+ "choice": opts[k], "adaptation": "native" if native else f"{strategy}", "protocol": PROTOCOL, "model": S["model"],
346
+ "usage": {"prompt_tokens": ctx.prompt_tokens, "completion_tokens": ctx.completion_tokens,
347
+ "total_tokens": ctx.prompt_tokens + ctx.completion_tokens},
348
+ "elapsed_seconds": time.time() - t0, "num_model_requests": ctx.requests}
349
+ if think is not None:
350
+ resp["thinking"] = {"mode": req.thinking, "threshold": tau, **think}
351
+ return resp
352
+
353
+
354
+ _build_app = entry.build_app
355
+
356
+
357
+ def build_app(args, *a, **kw):
358
+ app = _build_app(args, *a, **kw)
359
+ setup(args)
360
+ app.include_router(router)
361
+ return app
362
+
363
+
364
+ entry.build_app = build_app
365
+
366
+ if __name__ == "__main__":
367
+ entry.main()
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cc8d3a0ce36466ccc1278bf987df5f71db1719b9ca6b4118264f45cb627bfe0f
3
+ size 32169626
tokenizer_config.json ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "audio_token": "<|audio|>",
3
+ "backend": "tokenizers",
4
+ "boa_token": "<|audio>",
5
+ "boi_token": "<|image>",
6
+ "bos_token": "<bos>",
7
+ "eoa_token": "<audio|>",
8
+ "eoc_token": "<channel|>",
9
+ "eoi_token": "<image|>",
10
+ "eos_token": "<eos>",
11
+ "eot_token": "<turn|>",
12
+ "escape_token": "<|\"|>",
13
+ "etc_token": "<tool_call|>",
14
+ "etd_token": "<tool|>",
15
+ "etr_token": "<tool_response|>",
16
+ "extra_special_tokens": [
17
+ "<|video|>"
18
+ ],
19
+ "image_token": "<|image|>",
20
+ "mask_token": "<mask>",
21
+ "model_max_length": 1000000000000000019884624838656,
22
+ "pad_token": "<pad>",
23
+ "padding_side": "left",
24
+ "processor_class": "Gemma4Processor",
25
+ "response_schema": {
26
+ "type": "object",
27
+ "properties": {
28
+ "role": {
29
+ "const": "assistant"
30
+ },
31
+ "thinking": {
32
+ "type": "string"
33
+ },
34
+ "content": {
35
+ "type": "string"
36
+ },
37
+ "tool_calls": {
38
+ "x-regex-iterator": "<\\|tool_call>(.*?)<tool_call\\|>",
39
+ "type": "array",
40
+ "items": {
41
+ "type": "object",
42
+ "properties": {
43
+ "type": {
44
+ "const": "function"
45
+ },
46
+ "function": {
47
+ "type": "object",
48
+ "x-regex": "call\\:(?P<name>\\w+)(?P<arguments>\\{.*\\})",
49
+ "properties": {
50
+ "name": {
51
+ "type": "string"
52
+ },
53
+ "arguments": {
54
+ "type": "object",
55
+ "x-parser": "gemma4-tool-call",
56
+ "additionalProperties": {}
57
+ }
58
+ }
59
+ }
60
+ }
61
+ }
62
+ }
63
+ },
64
+ "x-regex": "(\\<\\|channel\\>thought\\n(?P<thinking>.*?)\\<channel\\|\\>)?(?P<tool_calls>\\<\\|tool_call\\>.*\\<tool_call\\|\\>)?(?P<content>(?:(?!\\<turn\\|\\>)(?!\\<\\|tool_response\\>).)+)?(?:\\<turn\\|\\>|\\<\\|tool_response\\>)?"
65
+ },
66
+ "response_template": {
67
+ "defaults": {
68
+ "role": "assistant"
69
+ },
70
+ "fields": {
71
+ "content": {
72
+ "close": [
73
+ "<turn|>",
74
+ "<|tool_response>",
75
+ "<eos>"
76
+ ],
77
+ "content": "text"
78
+ },
79
+ "thinking": {
80
+ "close": "<channel|>",
81
+ "content": "text",
82
+ "open": "<|channel>thought\n"
83
+ },
84
+ "tool_calls": {
85
+ "close": "<tool_call|>",
86
+ "content": "json",
87
+ "content_args": {
88
+ "string_delims": [
89
+ [
90
+ "<|\"|>",
91
+ "<|\"|>"
92
+ ]
93
+ ],
94
+ "unquoted_keys": true
95
+ },
96
+ "open_pattern": "<\\|tool_call>call:(?P<name>\\w+)",
97
+ "repeats": true,
98
+ "transform": {
99
+ "function": {
100
+ "arguments": "{content}",
101
+ "name": "{name}"
102
+ },
103
+ "type": "function"
104
+ }
105
+ }
106
+ },
107
+ "start_anchor": [
108
+ "<|turn>model\n",
109
+ "<tool_response|>"
110
+ ]
111
+ },
112
+ "soc_token": "<|channel>",
113
+ "sot_token": "<|turn>",
114
+ "stc_token": "<|tool_call>",
115
+ "std_token": "<|tool>",
116
+ "str_token": "<|tool_response>",
117
+ "think_token": "<|think|>",
118
+ "tokenizer_class": "GemmaTokenizer",
119
+ "unk_token": "<unk>"
120
+ }