RealFalconsAI commited on
Commit
99347a4
·
verified ·
1 Parent(s): c666ffc

Upload 16 files

Browse files
README.md ADDED
@@ -0,0 +1,243 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language: [en]
4
+ library_name: transformers
5
+ tags: [decision-model, system-one, calibrated-decisions, multiple-choice, zero-shot-classification, falcondec, lightdec]
6
+ ---
7
+ # LightDec_V2_Long v1.0.0 — FalconDec architecture
8
+
9
+ Single-pass, typed (`choice` / `noul` / `score`), calibrated closed-set decision model. Successor to
10
+ [`Falconsai/proof_v2`](https://huggingface.co/Falconsai/proof_v2).
11
+
12
+ * Built with the FalconDec notebook V3.7 · backbone `jhu-clsp/ettin-encoder-150m` · mode `scratch` · preset `long` · lineage: jhu-clsp/ettin-encoder-150m
13
+ * Parameters: 159.7M · weights: fp16 305 MB, int8 153 MB
14
+ * Layout: `[CLS] question [SEP] [MASK] opt1 … [MASK] optk [SEP] state [SEP]`; set-transformer option head;
15
+ temperature per (question type × option-count bucket).
16
+
17
+ ## Usage
18
+
19
+ ```python
20
+ import importlib.util
21
+ from huggingface_hub import snapshot_download
22
+ path = snapshot_download("<repo>") # or a local LightDec_V2_Long-v1.0.0 directory
23
+ spec = importlib.util.spec_from_file_location("falcondec_modeling", f"{path}/falcondec_modeling.py")
24
+ fdm = importlib.util.module_from_spec(spec); spec.loader.exec_module(fdm)
25
+ model, tok = fdm.load_falcondec(path) # add "/compact-int8" for the int8 artefact
26
+ out = fdm.decide(model, tok, state={"body": "I was charged twice, refund me or I cancel."}, questions={
27
+ "dept": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "payments", "tech": "bugs"}},
28
+ "churn": {"type": "noul", "instructions": "Does the user threaten to leave?"},
29
+ "urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["low", "medium", "high"]}})
30
+ ```
31
+
32
+ ## Test results
33
+
34
+ Overall: micro acc **0.784**, task-macro acc **0.789**, ECE **0.029**,
35
+ NLL 0.563, Brier 0.295, AURC 0.068.
36
+
37
+ | task | domain | heldout | n | chance | acc | proof_v2 (card) | ece |
38
+ |---|---|---|---|---|---|---|---|
39
+ | counsel/critique_quality | agentic | False | 201 | 0.333 | 0.627 | | 0.271 |
40
+ | agenttrek/finish_now | agentic | False | 151 | 0.500 | 0.788 | | 0.075 |
41
+ | counsel/step_has_error | agentic | False | 201 | 0.500 | 0.831 | | 0.148 |
42
+ | agenttrek/next_action_type | agentic | False | 487 | 0.191 | 0.864 | | 0.092 |
43
+ | hotpotqa/retrieve | agentic | False | 497 | 0.167 | 0.893 | | 0.043 |
44
+ | hotpotqa/comparison_yes_no | agentic | False | 26 | 0.500 | 0.923 | | 0.082 |
45
+ | sst5/score | classification | True | 500 | 0.200 | 0.406 | | 0.069 |
46
+ | emotion/6way | classification | True | 500 | 0.167 | 0.472 | | 0.211 |
47
+ | yelp/score | classification | False | 500 | 0.200 | 0.668 | | 0.091 |
48
+ | ag_news/topic | classification | False | 500 | 0.250 | 0.900 | | 0.052 |
49
+ | devign/vulnerability | code | False | 500 | 0.500 | 0.632 | 0.542 | 0.070 |
50
+ | humaneval/completion | code | True | 119 | 0.394 | 0.849 | 0.575 | 0.140 |
51
+ | mbpp/bugspot | code | False | 256 | 0.406 | 0.855 | 0.475 | 0.037 |
52
+ | codexglue/func_name | code | False | 467 | 0.255 | 0.955 | 0.901 | 0.039 |
53
+ | bigclonebench/clone | code | False | 500 | 0.500 | 0.962 | 0.383 | 0.031 |
54
+ | codexglue/doc_to_code | code | False | 504 | 0.279 | 0.976 | 0.961 | 0.013 |
55
+ | mbpp/solution | code | False | 500 | 0.250 | 0.978 | 0.992 | 0.020 |
56
+ | codexglue/code_to_doc | code | False | 504 | 0.262 | 0.990 | 0.969 | 0.010 |
57
+ | codexglue/lang_id | code | False | 504 | 0.235 | 1.000 | 0.997 | 0.001 |
58
+ | prompt_injections/detect | guardrails | True | 116 | 0.500 | 0.595 | | 0.333 |
59
+ | agentharm/refuse | guardrails | True | 416 | 0.500 | 0.700 | | 0.089 |
60
+ | civil_comments/toxic | guardrails | False | 500 | 0.500 | 0.926 | | 0.051 |
61
+ | jailbreak/detect | guardrails | False | 262 | 0.500 | 0.973 | | 0.031 |
62
+ | banking77/intent_77 | intents | True | 300 | 0.013 | 0.580 | | 0.241 |
63
+ | banking77/intent | intents | True | 500 | 0.321 | 0.922 | 0.883 | 0.037 |
64
+ | massive_en/intent | intents | False | 500 | 0.130 | 0.930 | | 0.028 |
65
+ | clinc150/intent | intents | False | 500 | 0.124 | 0.970 | 0.850 | 0.015 |
66
+ | long/contract_clause | long_context | False | 150 | 0.200 | 1.000 | | 0.000 |
67
+ | long/email_thread | long_context | False | 150 | 0.250 | 1.000 | | 0.000 |
68
+ | long/service_log | long_context | False | 150 | 0.200 | 1.000 | | 0.001 |
69
+ | policy/table_count_transfer | policy | False | 500 | 0.121 | 0.324 | | 0.541 |
70
+ | policy/invoice_total_transfer | policy | False | 500 | 0.500 | 0.474 | | 0.082 |
71
+ | policy/sla_urgency_transfer | policy | False | 500 | 0.250 | 0.606 | | 0.131 |
72
+ | policy/invoice_overdue_transfer | policy | False | 500 | 0.500 | 0.824 | | 0.020 |
73
+ | policy/count_threshold_transfer | policy | False | 500 | 0.105 | 0.880 | | 0.048 |
74
+ | policy/table_compare_transfer | policy | False | 500 | 0.500 | 0.928 | | 0.049 |
75
+ | policy/table_extreme_transfer | policy | False | 500 | 0.122 | 0.932 | | 0.053 |
76
+ | policy/free_shipping_transfer | policy | False | 500 | 0.500 | 0.938 | | 0.041 |
77
+ | policy/refund_approval_transfer | policy | False | 500 | 0.333 | 0.962 | | 0.019 |
78
+ | policy/access_control_transfer | policy | False | 500 | 0.333 | 1.000 | | 0.002 |
79
+ | policy/return_window_transfer | policy | False | 500 | 0.333 | 1.000 | | 0.026 |
80
+ | aqua_rat/math | reasoning | False | 247 | 0.200 | 0.340 | | 0.080 |
81
+ | mmlu/mcq | reasoning | True | 500 | 0.250 | 0.390 | | 0.155 |
82
+ | anli/nli | reasoning | False | 498 | 0.333 | 0.486 | | 0.146 |
83
+ | arc_challenge/mcq | reasoning | True | 500 | 0.250 | 0.490 | 0.308 | 0.168 |
84
+ | openbookqa/mcq | reasoning | False | 500 | 0.250 | 0.572 | 0.292 | 0.218 |
85
+ | hellaswag/continuation | reasoning | False | 500 | 0.250 | 0.576 | | 0.075 |
86
+ | arc_easy/mcq | reasoning | True | 500 | 0.250 | 0.612 | 0.425 | 0.122 |
87
+ | commonsense_qa/mcq | reasoning | False | 493 | 0.200 | 0.643 | 0.442 | 0.143 |
88
+ | winogrande/blank | reasoning | False | 500 | 0.500 | 0.662 | | 0.116 |
89
+ | gsm8k/math | reasoning | False | 500 | 0.250 | 0.700 | 0.275 | 0.085 |
90
+ | boolq/yes_no | reasoning | False | 500 | 0.500 | 0.822 | 0.717 | 0.080 |
91
+ | mnli/claim | reasoning | False | 500 | 0.333 | 0.860 | 0.492 | 0.081 |
92
+ | snli/nli | reasoning | False | 988 | 0.333 | 0.895 | | 0.108 |
93
+ | sciq/mcq | reasoning | False | 498 | 0.250 | 0.954 | 0.692 | 0.021 |
94
+ | scitail/support | reasoning | False | 500 | 0.500 | 0.962 | | 0.030 |
95
+ | qasc/mcq | reasoning | False | 500 | 0.125 | 0.986 | | 0.008 |
96
+ | snli/must_be_true | reasoning | False | 500 | 0.333 | 0.986 | 0.908 | 0.043 |
97
+ | snli/contradicts | reasoning | False | 500 | 0.333 | 0.990 | 0.892 | 0.019 |
98
+ | triage/support_email | support | False | 505 | 0.323 | 0.945 | | 0.046 |
99
+ | bitext/category | support | False | 500 | 0.161 | 1.000 | | 0.001 |
100
+ | bitext/route | support | False | 500 | 0.190 | 1.000 | 0.958 | 0.000 |
101
+ | tev1_test/sst5 | tev1_benchmark | True | 150 | 0.200 | 0.393 | | 0.108 |
102
+ | tev1_test/routing | tev1_benchmark | True | 600 | 0.200 | 0.458 | | 0.038 |
103
+ | tev1_test/policy | tev1_benchmark | True | 1200 | 0.333 | 0.529 | | 0.059 |
104
+ | tev1_test/banking77 | tev1_benchmark | True | 200 | 0.178 | 0.720 | | 0.137 |
105
+ | tev1_test/mnli | tev1_benchmark | True | 300 | 0.333 | 0.723 | | 0.075 |
106
+ | tev1_test/boolq | tev1_benchmark | True | 200 | 0.500 | 0.835 | | 0.050 |
107
+ | tev1_test/ag_news | tev1_benchmark | True | 150 | 0.250 | 0.920 | | 0.098 |
108
+ | typed_decisions/agent_trace_observability | workflows | False | 500 | 0.300 | 0.734 | | 0.222 |
109
+ | typed_decisions/customer_service | workflows | False | 500 | 0.280 | 0.764 | | 0.210 |
110
+ | typed_decisions/security_incidents | workflows | False | 500 | 0.340 | 0.768 | | 0.235 |
111
+ | typed_decisions/invoice_processing | workflows | False | 500 | 0.350 | 0.826 | | 0.195 |
112
+
113
+
114
+ ### Head-to-head with proof_v2 on identical test decisions
115
+
116
+ | task | n | LightDec | proof_v2 | Δ |
117
+ |---|---|---|---|---|
118
+ | ag_news/topic | 120.000 | 0.850 | 0.450 | 0.400 |
119
+ | agentharm/refuse | 120.000 | 0.667 | 0.450 | 0.217 |
120
+ | agenttrek/finish_now | 120.000 | 0.783 | 0.258 | 0.525 |
121
+ | agenttrek/next_action_type | 120.000 | 0.875 | 0.075 | 0.800 |
122
+ | anli/nli | 120.000 | 0.592 | 0.325 | 0.267 |
123
+ | aqua_rat/math | 120.000 | 0.350 | 0.242 | 0.108 |
124
+ | arc_challenge/mcq | 120.000 | 0.533 | 0.308 | 0.225 |
125
+ | arc_easy/mcq | 120.000 | 0.650 | 0.458 | 0.192 |
126
+ | banking77/intent | 120.000 | 0.925 | 0.867 | 0.058 |
127
+ | bigclonebench/clone | 120.000 | 0.967 | 0.242 | 0.725 |
128
+ | bitext/category | 120.000 | 1.000 | 0.858 | 0.142 |
129
+ | bitext/route | 120.000 | 1.000 | 0.925 | 0.075 |
130
+ | boolq/yes_no | 120.000 | 0.825 | 0.667 | 0.158 |
131
+ | civil_comments/toxic | 120.000 | 0.933 | 0.358 | 0.575 |
132
+ | clinc150/intent | 120.000 | 0.975 | 0.642 | 0.333 |
133
+ | codexglue/code_to_doc | 120.000 | 1.000 | 0.933 | 0.067 |
134
+ | codexglue/doc_to_code | 120.000 | 0.983 | 0.975 | 0.008 |
135
+ | codexglue/func_name | 120.000 | 0.992 | 0.917 | 0.075 |
136
+ | codexglue/lang_id | 120.000 | 1.000 | 1.000 | 0.000 |
137
+ | commonsense_qa/mcq | 120.000 | 0.683 | 0.408 | 0.275 |
138
+ | counsel/critique_quality | 120.000 | 0.608 | 0.258 | 0.350 |
139
+ | counsel/step_has_error | 120.000 | 0.825 | 0.750 | 0.075 |
140
+ | devign/vulnerability | 120.000 | 0.667 | 0.567 | 0.100 |
141
+ | emotion/6way | 120.000 | 0.492 | 0.367 | 0.125 |
142
+ | gsm8k/math | 120.000 | 0.733 | 0.125 | 0.608 |
143
+ | hellaswag/continuation | 120.000 | 0.608 | 0.292 | 0.317 |
144
+ | hotpotqa/comparison_yes_no | 26.000 | 0.923 | 0.462 | 0.462 |
145
+ | hotpotqa/retrieve | 120.000 | 0.875 | 0.292 | 0.583 |
146
+ | humaneval/completion | 119.000 | 0.849 | 0.513 | 0.336 |
147
+ | jailbreak/detect | 120.000 | 0.983 | 0.567 | 0.417 |
148
+ | long/contract_clause | 120.000 | 1.000 | 0.233 | 0.767 |
149
+ | long/email_thread | 120.000 | 1.000 | 0.050 | 0.950 |
150
+ | long/service_log | 120.000 | 1.000 | 0.242 | 0.758 |
151
+ | massive_en/intent | 120.000 | 0.950 | 0.742 | 0.208 |
152
+ | mbpp/bugspot | 120.000 | 0.858 | 0.575 | 0.283 |
153
+ | mbpp/solution | 120.000 | 1.000 | 0.850 | 0.150 |
154
+ | mmlu/mcq | 120.000 | 0.408 | 0.308 | 0.100 |
155
+ | mnli/claim | 120.000 | 0.875 | 0.317 | 0.558 |
156
+ | openbookqa/mcq | 120.000 | 0.650 | 0.283 | 0.367 |
157
+ | policy/access_control_transfer | 120.000 | 1.000 | 0.233 | 0.767 |
158
+ | policy/count_threshold_transfer | 120.000 | 0.867 | 0.108 | 0.758 |
159
+ | policy/free_shipping_transfer | 120.000 | 0.925 | 0.708 | 0.217 |
160
+ | policy/invoice_overdue_transfer | 120.000 | 0.808 | 0.658 | 0.150 |
161
+ | policy/invoice_total_transfer | 120.000 | 0.450 | 0.525 | -0.075 |
162
+ | policy/refund_approval_transfer | 120.000 | 0.958 | 0.325 | 0.633 |
163
+ | policy/return_window_transfer | 120.000 | 1.000 | 0.342 | 0.658 |
164
+ | policy/sla_urgency_transfer | 120.000 | 0.517 | 0.442 | 0.075 |
165
+ | policy/table_compare_transfer | 120.000 | 0.917 | 0.492 | 0.425 |
166
+ | policy/table_count_transfer | 120.000 | 0.317 | 0.158 | 0.158 |
167
+ | policy/table_extreme_transfer | 120.000 | 0.925 | 0.142 | 0.783 |
168
+ | prompt_injections/detect | 116.000 | 0.595 | 0.543 | 0.052 |
169
+ | qasc/mcq | 120.000 | 0.992 | 0.858 | 0.133 |
170
+ | sciq/mcq | 120.000 | 0.967 | 0.867 | 0.100 |
171
+ | scitail/support | 120.000 | 0.958 | 0.425 | 0.533 |
172
+ | snli/contradicts | 120.000 | 1.000 | 0.358 | 0.642 |
173
+ | snli/must_be_true | 120.000 | 0.992 | 0.925 | 0.067 |
174
+ | snli/nli | 120.000 | 0.908 | 0.342 | 0.567 |
175
+ | sst5/score | 120.000 | 0.383 | 0.283 | 0.100 |
176
+ | tev1_test/ag_news | 120.000 | 0.925 | 0.425 | 0.500 |
177
+ | tev1_test/banking77 | 120.000 | 0.717 | 0.550 | 0.167 |
178
+ | tev1_test/boolq | 120.000 | 0.842 | 0.558 | 0.283 |
179
+ | tev1_test/mnli | 120.000 | 0.750 | 0.392 | 0.358 |
180
+ | tev1_test/policy | 120.000 | 0.500 | 0.408 | 0.092 |
181
+ | tev1_test/routing | 120.000 | 0.442 | 0.292 | 0.150 |
182
+ | tev1_test/sst5 | 120.000 | 0.392 | 0.275 | 0.117 |
183
+ | triage/support_email | 120.000 | 0.958 | 0.392 | 0.567 |
184
+ | typed_decisions/agent_trace_observability | 120.000 | 0.808 | 0.333 | 0.475 |
185
+ | typed_decisions/customer_service | 120.000 | 0.742 | 0.283 | 0.458 |
186
+ | typed_decisions/invoice_processing | 120.000 | 0.817 | 0.475 | 0.342 |
187
+ | typed_decisions/security_incidents | 120.000 | 0.758 | 0.500 | 0.258 |
188
+ | winogrande/blank | 120.000 | 0.692 | 0.558 | 0.133 |
189
+ | yelp/score | 120.000 | 0.625 | 0.325 | 0.300 |
190
+
191
+ ### Tev1 benchmark (tev1 v2.1 test records, held out)
192
+
193
+ | task | n | accuracy | source also in LightDec training |
194
+ |---|---|---|---|
195
+ | tev1_test/ag_news | 150 | 0.920 | yes |
196
+ | tev1_test/banking77 | 200 | 0.720 | no |
197
+ | tev1_test/boolq | 200 | 0.835 | yes |
198
+ | tev1_test/mnli | 300 | 0.723 | yes |
199
+ | tev1_test/policy | 1200 | 0.529 | no |
200
+ | tev1_test/routing | 600 | 0.458 | no |
201
+ | tev1_test/sst5 | 150 | 0.393 | no |
202
+
203
+ All categories 0.584 on 2800 decisions; excluding sources LightDec also trains on 0.518. Tev1-4B's published results (880/1,000 main decisions,
204
+ 300/300 policy transfer) are reused development benchmarks and may not be the same records.
205
+
206
+ ### Support-email triage (`triage/support_email`, 101 held-out emails)
207
+
208
+ Per question: topic 0.990 · refund 1.000 · breakage 0.990 · anger 0.921 · judgment 0.822. Routing-pile accuracy (person / engineering / feature log / reply): 0.941.
209
+
210
+ ### Long-state tasks at max_len 2048
211
+
212
+ long/contract_clause 1.000 (chance 0.20) · long/email_thread 1.000 (chance 0.25) · long/service_log 1.000 (chance 0.20)
213
+
214
+ ### Latency (measured on NVIDIA RTX PRO 6000 Blackwell Server Edition)
215
+
216
+ | device | weights | questions per call | p50 ms | p95 ms |
217
+ |---|---|---|---|---|
218
+ | cuda | fp16 | 1 | 10.1 | 10.4 |
219
+ | cuda | fp16 | 5 | 11.3 | 11.4 |
220
+ | cuda | fp16 | 10 | 12.6 | 12.7 |
221
+ | cpu fp32 | int8 file | 1 | 48.4 | 48.8 |
222
+
223
+ Batched throughput: 2589 decisions/s.
224
+
225
+ ## Training
226
+
227
+ Strictly proper objective (log + 0.5·spherical + 1.0·RPS for ordinal), soft teacher targets where available,
228
+ task-balanced sampling (α = 0.5), NOTA augmentation (8%), per-epoch option reshuffling,
229
+ AdamW + LLRD 0.9, EMA 0.999, 10 epoch(s), RLCD stage off/reverted.
230
+ Training time 291.7 min on NVIDIA RTX PRO 6000 Blackwell Server Edition.
231
+
232
+ ## Limitations
233
+
234
+ * Closed world: the model ranks the options you give it. Add "None of the above" when that is a valid answer (it was trained with it).
235
+ * Multi-step arithmetic (GSM8K / AQuA) is a System-2 task; route low-confidence answers to an LLM (`defer` flag, threshold 0.7).
236
+ * English only; code judgement (vulnerabilities, clones) is still weak; check calibration on your own traffic.
237
+
238
+ ## References
239
+
240
+ Warner et al. 2024 (ModernBERT) · Weller et al. 2025 (Ettin) · Gneiting & Raftery 2007 (strictly proper scoring rules) ·
241
+ Guo et al. 2017 (temperature scaling) · Geifman & El-Yaniv 2017 (selective classification) · Shao et al. 2024 (GRPO) ·
242
+ Williams 1992 (REINFORCE) · Zaheer et al. 2017 / Lee et al. 2019 (Deep Sets / Set Transformer) · Hinton et al. 2015 (distillation) ·
243
+ Laya (convaiinnovations) · TypeSafe Jev · Together Tev1.
compact-int8/README.md ADDED
@@ -0,0 +1,243 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language: [en]
4
+ library_name: transformers
5
+ tags: [decision-model, system-one, calibrated-decisions, multiple-choice, zero-shot-classification, falcondec, lightdec]
6
+ ---
7
+ # LightDec_V2_Long v1.0.0 — FalconDec architecture
8
+
9
+ Single-pass, typed (`choice` / `noul` / `score`), calibrated closed-set decision model. Successor to
10
+ [`Falconsai/proof_v2`](https://huggingface.co/Falconsai/proof_v2).
11
+
12
+ * Built with the FalconDec notebook V3.7 · backbone `jhu-clsp/ettin-encoder-150m` · mode `scratch` · preset `long` · lineage: jhu-clsp/ettin-encoder-150m
13
+ * Parameters: 159.7M · weights: fp16 305 MB, int8 153 MB
14
+ * Layout: `[CLS] question [SEP] [MASK] opt1 … [MASK] optk [SEP] state [SEP]`; set-transformer option head;
15
+ temperature per (question type × option-count bucket).
16
+
17
+ ## Usage
18
+
19
+ ```python
20
+ import importlib.util
21
+ from huggingface_hub import snapshot_download
22
+ path = snapshot_download("<repo>") # or a local LightDec_V2_Long-v1.0.0 directory
23
+ spec = importlib.util.spec_from_file_location("falcondec_modeling", f"{path}/falcondec_modeling.py")
24
+ fdm = importlib.util.module_from_spec(spec); spec.loader.exec_module(fdm)
25
+ model, tok = fdm.load_falcondec(path) # add "/compact-int8" for the int8 artefact
26
+ out = fdm.decide(model, tok, state={"body": "I was charged twice, refund me or I cancel."}, questions={
27
+ "dept": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "payments", "tech": "bugs"}},
28
+ "churn": {"type": "noul", "instructions": "Does the user threaten to leave?"},
29
+ "urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["low", "medium", "high"]}})
30
+ ```
31
+
32
+ ## Test results
33
+
34
+ Overall: micro acc **0.784**, task-macro acc **0.789**, ECE **0.029**,
35
+ NLL 0.563, Brier 0.295, AURC 0.068.
36
+
37
+ | task | domain | heldout | n | chance | acc | proof_v2 (card) | ece |
38
+ |---|---|---|---|---|---|---|---|
39
+ | counsel/critique_quality | agentic | False | 201 | 0.333 | 0.627 | | 0.271 |
40
+ | agenttrek/finish_now | agentic | False | 151 | 0.500 | 0.788 | | 0.075 |
41
+ | counsel/step_has_error | agentic | False | 201 | 0.500 | 0.831 | | 0.148 |
42
+ | agenttrek/next_action_type | agentic | False | 487 | 0.191 | 0.864 | | 0.092 |
43
+ | hotpotqa/retrieve | agentic | False | 497 | 0.167 | 0.893 | | 0.043 |
44
+ | hotpotqa/comparison_yes_no | agentic | False | 26 | 0.500 | 0.923 | | 0.082 |
45
+ | sst5/score | classification | True | 500 | 0.200 | 0.406 | | 0.069 |
46
+ | emotion/6way | classification | True | 500 | 0.167 | 0.472 | | 0.211 |
47
+ | yelp/score | classification | False | 500 | 0.200 | 0.668 | | 0.091 |
48
+ | ag_news/topic | classification | False | 500 | 0.250 | 0.900 | | 0.052 |
49
+ | devign/vulnerability | code | False | 500 | 0.500 | 0.632 | 0.542 | 0.070 |
50
+ | humaneval/completion | code | True | 119 | 0.394 | 0.849 | 0.575 | 0.140 |
51
+ | mbpp/bugspot | code | False | 256 | 0.406 | 0.855 | 0.475 | 0.037 |
52
+ | codexglue/func_name | code | False | 467 | 0.255 | 0.955 | 0.901 | 0.039 |
53
+ | bigclonebench/clone | code | False | 500 | 0.500 | 0.962 | 0.383 | 0.031 |
54
+ | codexglue/doc_to_code | code | False | 504 | 0.279 | 0.976 | 0.961 | 0.013 |
55
+ | mbpp/solution | code | False | 500 | 0.250 | 0.978 | 0.992 | 0.020 |
56
+ | codexglue/code_to_doc | code | False | 504 | 0.262 | 0.990 | 0.969 | 0.010 |
57
+ | codexglue/lang_id | code | False | 504 | 0.235 | 1.000 | 0.997 | 0.001 |
58
+ | prompt_injections/detect | guardrails | True | 116 | 0.500 | 0.595 | | 0.333 |
59
+ | agentharm/refuse | guardrails | True | 416 | 0.500 | 0.700 | | 0.089 |
60
+ | civil_comments/toxic | guardrails | False | 500 | 0.500 | 0.926 | | 0.051 |
61
+ | jailbreak/detect | guardrails | False | 262 | 0.500 | 0.973 | | 0.031 |
62
+ | banking77/intent_77 | intents | True | 300 | 0.013 | 0.580 | | 0.241 |
63
+ | banking77/intent | intents | True | 500 | 0.321 | 0.922 | 0.883 | 0.037 |
64
+ | massive_en/intent | intents | False | 500 | 0.130 | 0.930 | | 0.028 |
65
+ | clinc150/intent | intents | False | 500 | 0.124 | 0.970 | 0.850 | 0.015 |
66
+ | long/contract_clause | long_context | False | 150 | 0.200 | 1.000 | | 0.000 |
67
+ | long/email_thread | long_context | False | 150 | 0.250 | 1.000 | | 0.000 |
68
+ | long/service_log | long_context | False | 150 | 0.200 | 1.000 | | 0.001 |
69
+ | policy/table_count_transfer | policy | False | 500 | 0.121 | 0.324 | | 0.541 |
70
+ | policy/invoice_total_transfer | policy | False | 500 | 0.500 | 0.474 | | 0.082 |
71
+ | policy/sla_urgency_transfer | policy | False | 500 | 0.250 | 0.606 | | 0.131 |
72
+ | policy/invoice_overdue_transfer | policy | False | 500 | 0.500 | 0.824 | | 0.020 |
73
+ | policy/count_threshold_transfer | policy | False | 500 | 0.105 | 0.880 | | 0.048 |
74
+ | policy/table_compare_transfer | policy | False | 500 | 0.500 | 0.928 | | 0.049 |
75
+ | policy/table_extreme_transfer | policy | False | 500 | 0.122 | 0.932 | | 0.053 |
76
+ | policy/free_shipping_transfer | policy | False | 500 | 0.500 | 0.938 | | 0.041 |
77
+ | policy/refund_approval_transfer | policy | False | 500 | 0.333 | 0.962 | | 0.019 |
78
+ | policy/access_control_transfer | policy | False | 500 | 0.333 | 1.000 | | 0.002 |
79
+ | policy/return_window_transfer | policy | False | 500 | 0.333 | 1.000 | | 0.026 |
80
+ | aqua_rat/math | reasoning | False | 247 | 0.200 | 0.340 | | 0.080 |
81
+ | mmlu/mcq | reasoning | True | 500 | 0.250 | 0.390 | | 0.155 |
82
+ | anli/nli | reasoning | False | 498 | 0.333 | 0.486 | | 0.146 |
83
+ | arc_challenge/mcq | reasoning | True | 500 | 0.250 | 0.490 | 0.308 | 0.168 |
84
+ | openbookqa/mcq | reasoning | False | 500 | 0.250 | 0.572 | 0.292 | 0.218 |
85
+ | hellaswag/continuation | reasoning | False | 500 | 0.250 | 0.576 | | 0.075 |
86
+ | arc_easy/mcq | reasoning | True | 500 | 0.250 | 0.612 | 0.425 | 0.122 |
87
+ | commonsense_qa/mcq | reasoning | False | 493 | 0.200 | 0.643 | 0.442 | 0.143 |
88
+ | winogrande/blank | reasoning | False | 500 | 0.500 | 0.662 | | 0.116 |
89
+ | gsm8k/math | reasoning | False | 500 | 0.250 | 0.700 | 0.275 | 0.085 |
90
+ | boolq/yes_no | reasoning | False | 500 | 0.500 | 0.822 | 0.717 | 0.080 |
91
+ | mnli/claim | reasoning | False | 500 | 0.333 | 0.860 | 0.492 | 0.081 |
92
+ | snli/nli | reasoning | False | 988 | 0.333 | 0.895 | | 0.108 |
93
+ | sciq/mcq | reasoning | False | 498 | 0.250 | 0.954 | 0.692 | 0.021 |
94
+ | scitail/support | reasoning | False | 500 | 0.500 | 0.962 | | 0.030 |
95
+ | qasc/mcq | reasoning | False | 500 | 0.125 | 0.986 | | 0.008 |
96
+ | snli/must_be_true | reasoning | False | 500 | 0.333 | 0.986 | 0.908 | 0.043 |
97
+ | snli/contradicts | reasoning | False | 500 | 0.333 | 0.990 | 0.892 | 0.019 |
98
+ | triage/support_email | support | False | 505 | 0.323 | 0.945 | | 0.046 |
99
+ | bitext/category | support | False | 500 | 0.161 | 1.000 | | 0.001 |
100
+ | bitext/route | support | False | 500 | 0.190 | 1.000 | 0.958 | 0.000 |
101
+ | tev1_test/sst5 | tev1_benchmark | True | 150 | 0.200 | 0.393 | | 0.108 |
102
+ | tev1_test/routing | tev1_benchmark | True | 600 | 0.200 | 0.458 | | 0.038 |
103
+ | tev1_test/policy | tev1_benchmark | True | 1200 | 0.333 | 0.529 | | 0.059 |
104
+ | tev1_test/banking77 | tev1_benchmark | True | 200 | 0.178 | 0.720 | | 0.137 |
105
+ | tev1_test/mnli | tev1_benchmark | True | 300 | 0.333 | 0.723 | | 0.075 |
106
+ | tev1_test/boolq | tev1_benchmark | True | 200 | 0.500 | 0.835 | | 0.050 |
107
+ | tev1_test/ag_news | tev1_benchmark | True | 150 | 0.250 | 0.920 | | 0.098 |
108
+ | typed_decisions/agent_trace_observability | workflows | False | 500 | 0.300 | 0.734 | | 0.222 |
109
+ | typed_decisions/customer_service | workflows | False | 500 | 0.280 | 0.764 | | 0.210 |
110
+ | typed_decisions/security_incidents | workflows | False | 500 | 0.340 | 0.768 | | 0.235 |
111
+ | typed_decisions/invoice_processing | workflows | False | 500 | 0.350 | 0.826 | | 0.195 |
112
+
113
+
114
+ ### Head-to-head with proof_v2 on identical test decisions
115
+
116
+ | task | n | LightDec | proof_v2 | Δ |
117
+ |---|---|---|---|---|
118
+ | ag_news/topic | 120.000 | 0.850 | 0.450 | 0.400 |
119
+ | agentharm/refuse | 120.000 | 0.667 | 0.450 | 0.217 |
120
+ | agenttrek/finish_now | 120.000 | 0.783 | 0.258 | 0.525 |
121
+ | agenttrek/next_action_type | 120.000 | 0.875 | 0.075 | 0.800 |
122
+ | anli/nli | 120.000 | 0.592 | 0.325 | 0.267 |
123
+ | aqua_rat/math | 120.000 | 0.350 | 0.242 | 0.108 |
124
+ | arc_challenge/mcq | 120.000 | 0.533 | 0.308 | 0.225 |
125
+ | arc_easy/mcq | 120.000 | 0.650 | 0.458 | 0.192 |
126
+ | banking77/intent | 120.000 | 0.925 | 0.867 | 0.058 |
127
+ | bigclonebench/clone | 120.000 | 0.967 | 0.242 | 0.725 |
128
+ | bitext/category | 120.000 | 1.000 | 0.858 | 0.142 |
129
+ | bitext/route | 120.000 | 1.000 | 0.925 | 0.075 |
130
+ | boolq/yes_no | 120.000 | 0.825 | 0.667 | 0.158 |
131
+ | civil_comments/toxic | 120.000 | 0.933 | 0.358 | 0.575 |
132
+ | clinc150/intent | 120.000 | 0.975 | 0.642 | 0.333 |
133
+ | codexglue/code_to_doc | 120.000 | 1.000 | 0.933 | 0.067 |
134
+ | codexglue/doc_to_code | 120.000 | 0.983 | 0.975 | 0.008 |
135
+ | codexglue/func_name | 120.000 | 0.992 | 0.917 | 0.075 |
136
+ | codexglue/lang_id | 120.000 | 1.000 | 1.000 | 0.000 |
137
+ | commonsense_qa/mcq | 120.000 | 0.683 | 0.408 | 0.275 |
138
+ | counsel/critique_quality | 120.000 | 0.608 | 0.258 | 0.350 |
139
+ | counsel/step_has_error | 120.000 | 0.825 | 0.750 | 0.075 |
140
+ | devign/vulnerability | 120.000 | 0.667 | 0.567 | 0.100 |
141
+ | emotion/6way | 120.000 | 0.492 | 0.367 | 0.125 |
142
+ | gsm8k/math | 120.000 | 0.733 | 0.125 | 0.608 |
143
+ | hellaswag/continuation | 120.000 | 0.608 | 0.292 | 0.317 |
144
+ | hotpotqa/comparison_yes_no | 26.000 | 0.923 | 0.462 | 0.462 |
145
+ | hotpotqa/retrieve | 120.000 | 0.875 | 0.292 | 0.583 |
146
+ | humaneval/completion | 119.000 | 0.849 | 0.513 | 0.336 |
147
+ | jailbreak/detect | 120.000 | 0.983 | 0.567 | 0.417 |
148
+ | long/contract_clause | 120.000 | 1.000 | 0.233 | 0.767 |
149
+ | long/email_thread | 120.000 | 1.000 | 0.050 | 0.950 |
150
+ | long/service_log | 120.000 | 1.000 | 0.242 | 0.758 |
151
+ | massive_en/intent | 120.000 | 0.950 | 0.742 | 0.208 |
152
+ | mbpp/bugspot | 120.000 | 0.858 | 0.575 | 0.283 |
153
+ | mbpp/solution | 120.000 | 1.000 | 0.850 | 0.150 |
154
+ | mmlu/mcq | 120.000 | 0.408 | 0.308 | 0.100 |
155
+ | mnli/claim | 120.000 | 0.875 | 0.317 | 0.558 |
156
+ | openbookqa/mcq | 120.000 | 0.650 | 0.283 | 0.367 |
157
+ | policy/access_control_transfer | 120.000 | 1.000 | 0.233 | 0.767 |
158
+ | policy/count_threshold_transfer | 120.000 | 0.867 | 0.108 | 0.758 |
159
+ | policy/free_shipping_transfer | 120.000 | 0.925 | 0.708 | 0.217 |
160
+ | policy/invoice_overdue_transfer | 120.000 | 0.808 | 0.658 | 0.150 |
161
+ | policy/invoice_total_transfer | 120.000 | 0.450 | 0.525 | -0.075 |
162
+ | policy/refund_approval_transfer | 120.000 | 0.958 | 0.325 | 0.633 |
163
+ | policy/return_window_transfer | 120.000 | 1.000 | 0.342 | 0.658 |
164
+ | policy/sla_urgency_transfer | 120.000 | 0.517 | 0.442 | 0.075 |
165
+ | policy/table_compare_transfer | 120.000 | 0.917 | 0.492 | 0.425 |
166
+ | policy/table_count_transfer | 120.000 | 0.317 | 0.158 | 0.158 |
167
+ | policy/table_extreme_transfer | 120.000 | 0.925 | 0.142 | 0.783 |
168
+ | prompt_injections/detect | 116.000 | 0.595 | 0.543 | 0.052 |
169
+ | qasc/mcq | 120.000 | 0.992 | 0.858 | 0.133 |
170
+ | sciq/mcq | 120.000 | 0.967 | 0.867 | 0.100 |
171
+ | scitail/support | 120.000 | 0.958 | 0.425 | 0.533 |
172
+ | snli/contradicts | 120.000 | 1.000 | 0.358 | 0.642 |
173
+ | snli/must_be_true | 120.000 | 0.992 | 0.925 | 0.067 |
174
+ | snli/nli | 120.000 | 0.908 | 0.342 | 0.567 |
175
+ | sst5/score | 120.000 | 0.383 | 0.283 | 0.100 |
176
+ | tev1_test/ag_news | 120.000 | 0.925 | 0.425 | 0.500 |
177
+ | tev1_test/banking77 | 120.000 | 0.717 | 0.550 | 0.167 |
178
+ | tev1_test/boolq | 120.000 | 0.842 | 0.558 | 0.283 |
179
+ | tev1_test/mnli | 120.000 | 0.750 | 0.392 | 0.358 |
180
+ | tev1_test/policy | 120.000 | 0.500 | 0.408 | 0.092 |
181
+ | tev1_test/routing | 120.000 | 0.442 | 0.292 | 0.150 |
182
+ | tev1_test/sst5 | 120.000 | 0.392 | 0.275 | 0.117 |
183
+ | triage/support_email | 120.000 | 0.958 | 0.392 | 0.567 |
184
+ | typed_decisions/agent_trace_observability | 120.000 | 0.808 | 0.333 | 0.475 |
185
+ | typed_decisions/customer_service | 120.000 | 0.742 | 0.283 | 0.458 |
186
+ | typed_decisions/invoice_processing | 120.000 | 0.817 | 0.475 | 0.342 |
187
+ | typed_decisions/security_incidents | 120.000 | 0.758 | 0.500 | 0.258 |
188
+ | winogrande/blank | 120.000 | 0.692 | 0.558 | 0.133 |
189
+ | yelp/score | 120.000 | 0.625 | 0.325 | 0.300 |
190
+
191
+ ### Tev1 benchmark (tev1 v2.1 test records, held out)
192
+
193
+ | task | n | accuracy | source also in LightDec training |
194
+ |---|---|---|---|
195
+ | tev1_test/ag_news | 150 | 0.920 | yes |
196
+ | tev1_test/banking77 | 200 | 0.720 | no |
197
+ | tev1_test/boolq | 200 | 0.835 | yes |
198
+ | tev1_test/mnli | 300 | 0.723 | yes |
199
+ | tev1_test/policy | 1200 | 0.529 | no |
200
+ | tev1_test/routing | 600 | 0.458 | no |
201
+ | tev1_test/sst5 | 150 | 0.393 | no |
202
+
203
+ All categories 0.584 on 2800 decisions; excluding sources LightDec also trains on 0.518. Tev1-4B's published results (880/1,000 main decisions,
204
+ 300/300 policy transfer) are reused development benchmarks and may not be the same records.
205
+
206
+ ### Support-email triage (`triage/support_email`, 101 held-out emails)
207
+
208
+ Per question: topic 0.990 · refund 1.000 · breakage 0.990 · anger 0.921 · judgment 0.822. Routing-pile accuracy (person / engineering / feature log / reply): 0.941.
209
+
210
+ ### Long-state tasks at max_len 2048
211
+
212
+ long/contract_clause 1.000 (chance 0.20) · long/email_thread 1.000 (chance 0.25) · long/service_log 1.000 (chance 0.20)
213
+
214
+ ### Latency (measured on NVIDIA RTX PRO 6000 Blackwell Server Edition)
215
+
216
+ | device | weights | questions per call | p50 ms | p95 ms |
217
+ |---|---|---|---|---|
218
+ | cuda | fp16 | 1 | 10.1 | 10.4 |
219
+ | cuda | fp16 | 5 | 11.3 | 11.4 |
220
+ | cuda | fp16 | 10 | 12.6 | 12.7 |
221
+ | cpu fp32 | int8 file | 1 | 48.4 | 48.8 |
222
+
223
+ Batched throughput: 2589 decisions/s.
224
+
225
+ ## Training
226
+
227
+ Strictly proper objective (log + 0.5·spherical + 1.0·RPS for ordinal), soft teacher targets where available,
228
+ task-balanced sampling (α = 0.5), NOTA augmentation (8%), per-epoch option reshuffling,
229
+ AdamW + LLRD 0.9, EMA 0.999, 10 epoch(s), RLCD stage off/reverted.
230
+ Training time 291.7 min on NVIDIA RTX PRO 6000 Blackwell Server Edition.
231
+
232
+ ## Limitations
233
+
234
+ * Closed world: the model ranks the options you give it. Add "None of the above" when that is a valid answer (it was trained with it).
235
+ * Multi-step arithmetic (GSM8K / AQuA) is a System-2 task; route low-confidence answers to an LLM (`defer` flag, threshold 0.7).
236
+ * English only; code judgement (vulnerabilities, clones) is still weak; check calibration on your own traffic.
237
+
238
+ ## References
239
+
240
+ Warner et al. 2024 (ModernBERT) · Weller et al. 2025 (Ettin) · Gneiting & Raftery 2007 (strictly proper scoring rules) ·
241
+ Guo et al. 2017 (temperature scaling) · Geifman & El-Yaniv 2017 (selective classification) · Shao et al. 2024 (GRPO) ·
242
+ Williams 1992 (REINFORCE) · Zaheer et al. 2017 / Lee et al. 2019 (Deep Sets / Set Transformer) · Hinton et al. 2015 (distillation) ·
243
+ Laya (convaiinnovations) · TypeSafe Jev · Together Tev1.
compact-int8/encoder/config.json ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "ModernBertForMaskedLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 50281,
8
+ "causal_mask": false,
9
+ "classifier_activation": "gelu",
10
+ "classifier_bias": false,
11
+ "classifier_dropout": 0.0,
12
+ "classifier_pooling": "mean",
13
+ "cls_token_id": 50281,
14
+ "decoder_bias": true,
15
+ "deterministic_flash_attn": false,
16
+ "dtype": "float32",
17
+ "embedding_dropout": 0.0,
18
+ "eos_token_id": 50282,
19
+ "global_attn_every_n_layers": 3,
20
+ "gradient_checkpointing": false,
21
+ "hidden_activation": "gelu",
22
+ "hidden_size": 768,
23
+ "initializer_cutoff_factor": 2.0,
24
+ "initializer_range": 0.02,
25
+ "intermediate_size": 1152,
26
+ "is_causal": false,
27
+ "layer_norm_eps": 1e-05,
28
+ "layer_types": [
29
+ "full_attention",
30
+ "sliding_attention",
31
+ "sliding_attention",
32
+ "full_attention",
33
+ "sliding_attention",
34
+ "sliding_attention",
35
+ "full_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "full_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "full_attention",
42
+ "sliding_attention",
43
+ "sliding_attention",
44
+ "full_attention",
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "full_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "full_attention"
51
+ ],
52
+ "local_attention": 128,
53
+ "max_position_embeddings": 7999,
54
+ "mlp_bias": false,
55
+ "mlp_dropout": 0.0,
56
+ "model_type": "modernbert",
57
+ "norm_bias": false,
58
+ "norm_eps": 1e-05,
59
+ "num_attention_heads": 12,
60
+ "num_hidden_layers": 22,
61
+ "pad_token_id": 50283,
62
+ "position_embedding_type": "sans_pos",
63
+ "rope_parameters": {
64
+ "full_attention": {
65
+ "rope_theta": 160000.0,
66
+ "rope_type": "default"
67
+ },
68
+ "sliding_attention": {
69
+ "rope_theta": 160000.0,
70
+ "rope_type": "default"
71
+ }
72
+ },
73
+ "sep_token_id": 50282,
74
+ "sparse_pred_ignore_index": -100,
75
+ "sparse_prediction": false,
76
+ "tie_word_embeddings": true,
77
+ "transformers_version": "5.17.0",
78
+ "vocab_size": 50368
79
+ }
compact-int8/falcondec_config.json ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "max_len": 2048,
3
+ "long_max_len": 2048,
4
+ "long_opts_threshold": 24,
5
+ "head_max_len": 192,
6
+ "max_tok_per_opt": 24,
7
+ "max_opts_single_pass": 96,
8
+ "interact_layers": 2,
9
+ "interact_heads": 8,
10
+ "name": "LightDec_V2_Long",
11
+ "backbone": "jhu-clsp/ettin-encoder-150m",
12
+ "backbone_init": "pretrained",
13
+ "lineage": [
14
+ "jhu-clsp/ettin-encoder-150m"
15
+ ],
16
+ "parent_version": null,
17
+ "special": {
18
+ "cls": 50281,
19
+ "sep": 50282,
20
+ "mask": 50284,
21
+ "pad": 50283
22
+ },
23
+ "defer_threshold": 0.7,
24
+ "qtypes": {
25
+ "choice": 0,
26
+ "noul": 1,
27
+ "score": 2
28
+ },
29
+ "n_buckets": 4,
30
+ "version": "1.0.0",
31
+ "notebook_version": "V3.7",
32
+ "created": "2026-09-28T10:52:36.426087Z",
33
+ "temperature": [
34
+ [
35
+ 2.37591814994812,
36
+ 2.0426506996154785,
37
+ 1.6455482244491577,
38
+ 1.8417285680770874
39
+ ],
40
+ [
41
+ 2.4885058403015137,
42
+ 2.4885058403015137,
43
+ 2.4885058403015137,
44
+ 2.4885058403015137
45
+ ],
46
+ [
47
+ 2.2925639152526855,
48
+ 2.2925639152526855,
49
+ 2.2925639152526855,
50
+ 2.2925639152526855
51
+ ]
52
+ ],
53
+ "weights": "model_int8.safetensors"
54
+ }
compact-int8/falcondec_modeling.py ADDED
@@ -0,0 +1,358 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """FalconDec — Falcon Decision model: single-pass, typed, calibrated closed-set decisions.
3
+
4
+ Layout : [CLS] question [SEP] [MASK] opt_1 ... [MASK] opt_k [SEP] state [SEP]
5
+ Head : marker vectors + CLS context + question-type embedding
6
+ -> permutation-equivariant set transformer (options attend to each other; no positions)
7
+ -> MLP -> one logit per option -> softmax over this question's options
8
+ Calib. : temperature per (question type, option-count bucket), stored in the checkpoint
9
+ """
10
+ from __future__ import annotations
11
+
12
+ import contextlib
13
+ import json
14
+ import shutil
15
+ from pathlib import Path
16
+
17
+ import numpy as np
18
+ import torch
19
+ import torch.nn as nn
20
+
21
+ QTYPES = {"choice": 0, "noul": 1, "score": 2}
22
+ N_BUCKETS = 4
23
+ CONFIG_FILE = "falcondec_config.json"
24
+ FP16_FILE = "model.safetensors"
25
+ INT8_FILE = "model_int8.safetensors"
26
+ MODEL_KEYS = ("input_ids", "attention_mask", "marker_pos", "marker_mask", "qtype")
27
+
28
+
29
+ def n_bucket(n: int) -> int:
30
+ return 0 if n <= 2 else 1 if n <= 5 else 2 if n <= 12 else 3
31
+
32
+
33
+ def special_ids(tok) -> dict:
34
+ sp = {"cls": tok.cls_token_id, "sep": tok.sep_token_id, "mask": tok.mask_token_id, "pad": tok.pad_token_id}
35
+ if sp["cls"] is None:
36
+ sp["cls"] = tok.bos_token_id
37
+ if sp["sep"] is None:
38
+ sp["sep"] = tok.eos_token_id
39
+ if sp["pad"] is None:
40
+ sp["pad"] = sp["sep"]
41
+ if sp["mask"] is None:
42
+ raise ValueError("FalconDec needs a tokenizer with a mask token (used as the option marker).")
43
+ return {k: int(v) for k, v in sp.items()}
44
+
45
+
46
+ def _as_list(x):
47
+ return x.tolist() if hasattr(x, "tolist") else list(x)
48
+
49
+
50
+ def assemble(q_ids, opt_ids, s_ids, sp, max_len=512, head_max_len=192, max_tok_per_opt=24,
51
+ long_max_len=2048, long_opts_threshold=24):
52
+ """Build one input sequence. Returns (ids, marker_positions)."""
53
+ n = len(opt_ids)
54
+ eff = max_len if n <= long_opts_threshold else max(max_len, long_max_len)
55
+ q = _as_list(q_ids)[:96]
56
+ need = len(q) + 2 + n * (max_tok_per_opt + 1)
57
+ head = min(eff - 64, max(head_max_len, need))
58
+ per = max(2, min(max_tok_per_opt, (head - len(q) - 2) // max(n, 1) - 1))
59
+ ids = [sp["cls"]] + q + [sp["sep"]]
60
+ markers = []
61
+ for o in opt_ids:
62
+ markers.append(len(ids))
63
+ ids.append(sp["mask"])
64
+ ids.extend(_as_list(o[:per]))
65
+ ids.append(sp["sep"])
66
+ room = eff - len(ids) - 1
67
+ if room > 0 and s_ids is not None and len(s_ids) > 0:
68
+ ids.extend(_as_list(s_ids[:room]))
69
+ ids.append(sp["sep"])
70
+ if len(ids) > eff:
71
+ ids = ids[:eff]
72
+ if markers and markers[-1] >= len(ids):
73
+ raise ValueError("Too many options for one pass; reduce options or raise long_max_len.")
74
+ return ids, markers
75
+
76
+
77
+ def collate_features(feats, pad_id, device=None):
78
+ """feats: list of (ids, markers, qtype_index) -> dict of padded tensors matching FalconDec.forward."""
79
+ B = len(feats)
80
+ T = max(len(f[0]) for f in feats)
81
+ K = max(len(f[1]) for f in feats)
82
+ ids = torch.full((B, T), pad_id, dtype=torch.long)
83
+ att = torch.zeros((B, T), dtype=torch.long)
84
+ mpos = torch.zeros((B, K), dtype=torch.long)
85
+ mmask = torch.zeros((B, K), dtype=torch.bool)
86
+ qt = torch.zeros(B, dtype=torch.long)
87
+ for i, (x, mk, q) in enumerate(feats):
88
+ ids[i, : len(x)] = torch.as_tensor(x, dtype=torch.long)
89
+ att[i, : len(x)] = 1
90
+ mpos[i, : len(mk)] = torch.as_tensor(mk, dtype=torch.long)
91
+ mmask[i, : len(mk)] = True
92
+ qt[i] = int(q)
93
+ out = dict(input_ids=ids, attention_mask=att, marker_pos=mpos, marker_mask=mmask, qtype=qt)
94
+ if device is not None:
95
+ out = {k: v.to(device, non_blocking=True) for k, v in out.items()}
96
+ return out
97
+
98
+
99
+ class OptionInteraction(nn.Module):
100
+ """Set transformer over the options of one question (no positional encoding => order-equivariant)."""
101
+
102
+ def __init__(self, d, n_layers=2, n_heads=8, dropout=0.1):
103
+ super().__init__()
104
+ layer = nn.TransformerEncoderLayer(d, n_heads, dim_feedforward=2 * d, dropout=dropout,
105
+ activation="gelu", batch_first=True, norm_first=True)
106
+ self.enc = nn.TransformerEncoder(layer, n_layers, enable_nested_tensor=False)
107
+
108
+ def forward(self, x, mask):
109
+ return self.enc(x, src_key_padding_mask=~mask)
110
+
111
+
112
+ class FalconDec(nn.Module):
113
+ def __init__(self, encoder, fcfg: dict):
114
+ super().__init__()
115
+ self.encoder = encoder
116
+ self.fcfg = dict(fcfg)
117
+ d = encoder.config.hidden_size
118
+ self.qtype_emb = nn.Embedding(len(QTYPES), d)
119
+ self.ctx_proj = nn.Linear(d, d)
120
+ self.opt_norm = nn.LayerNorm(d)
121
+ self.interact = OptionInteraction(d, int(self.fcfg.get("interact_layers", 2)),
122
+ int(self.fcfg.get("interact_heads", 8)))
123
+ self.scorer = nn.Sequential(nn.Linear(d, d), nn.GELU(), nn.Dropout(0.1), nn.Linear(d, 1))
124
+ self.register_buffer("temperature", torch.ones(len(QTYPES), N_BUCKETS))
125
+
126
+ @property
127
+ def device(self):
128
+ return next(self.parameters()).device
129
+
130
+ def num_parameters(self):
131
+ return sum(p.numel() for p in self.parameters())
132
+
133
+ def forward(self, input_ids, attention_mask, marker_pos, marker_mask, qtype):
134
+ h = self.encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state
135
+ B, K = marker_pos.shape
136
+ idx = marker_pos.unsqueeze(-1).expand(B, K, h.size(-1))
137
+ opt = h.gather(1, idx)
138
+ ctx = self.ctx_proj(h[:, 0]).unsqueeze(1)
139
+ x = self.opt_norm(opt + ctx + self.qtype_emb(qtype).unsqueeze(1))
140
+ x = self.interact(x, marker_mask)
141
+ logits = self.scorer(x).squeeze(-1).float()
142
+ return logits.masked_fill(~marker_mask, -1e4)
143
+
144
+
145
+ # ------------------------------------------------------------------ storage
146
+ def clean_state_dict(sd):
147
+ return {k.replace("_orig_mod.", ""): v.detach().cpu().contiguous() for k, v in sd.items()}
148
+
149
+
150
+ def quantize_int8(sd, min_numel=4096):
151
+ """Per-output-channel symmetric int8 for every matrix; fp16 for everything else."""
152
+ out = {}
153
+ for k, v in sd.items():
154
+ if v.is_floating_point() and v.ndim == 2 and v.numel() >= min_numel:
155
+ w = v.float()
156
+ s = (w.abs().amax(dim=1, keepdim=True) / 127.0).clamp_min(1e-12)
157
+ out[k] = torch.round(w / s).clamp_(-127, 127).to(torch.int8).contiguous()
158
+ out[k + "::scale"] = s.contiguous()
159
+ elif v.is_floating_point():
160
+ out[k] = v.half().contiguous()
161
+ else:
162
+ out[k] = v.contiguous()
163
+ return out
164
+
165
+
166
+ def dequantize_int8(sd):
167
+ out = {}
168
+ for k, v in sd.items():
169
+ if k.endswith("::scale"):
170
+ continue
171
+ s = sd.get(k + "::scale")
172
+ out[k] = (v.float() * s).half() if s is not None else v
173
+ return out
174
+
175
+
176
+ def save_falcondec(model, tok, out_dir, int8=False, extra_files=None):
177
+ from safetensors.torch import save_file
178
+ out = Path(out_dir)
179
+ out.mkdir(parents=True, exist_ok=True)
180
+ enc_cfg = model.encoder.config
181
+ if hasattr(enc_cfg, "reference_compile"):
182
+ enc_cfg.reference_compile = False
183
+ enc_cfg.save_pretrained(str(out / "encoder"))
184
+ tok.save_pretrained(str(out / "tokenizer"))
185
+ sd = clean_state_dict(model.state_dict())
186
+ fc = dict(model.fcfg)
187
+ fc["temperature"] = model.temperature.detach().float().cpu().tolist()
188
+ fc["weights"] = INT8_FILE if int8 else FP16_FILE
189
+ if int8:
190
+ save_file(quantize_int8(sd), str(out / INT8_FILE), metadata={"format": "falcondec-int8"})
191
+ else:
192
+ save_file({k: (v.half() if v.is_floating_point() else v) for k, v in sd.items()},
193
+ str(out / FP16_FILE), metadata={"format": "falcondec-fp16"})
194
+ (out / CONFIG_FILE).write_text(json.dumps(fc, indent=2, default=str), encoding="utf-8")
195
+ try:
196
+ shutil.copy(__file__, out / "falcondec_modeling.py")
197
+ except Exception:
198
+ pass
199
+ for name, content in (extra_files or {}).items():
200
+ (out / name).write_text(content, encoding="utf-8")
201
+ return out
202
+
203
+
204
+ def load_falcondec(path, device=None, dtype=None, attn_implementation="sdpa"):
205
+ """Load a FalconDec directory (fp16 or int8) or Hub repo. Returns (model, tokenizer)."""
206
+ from safetensors.torch import load_file
207
+ from transformers import AutoConfig, AutoModel, AutoTokenizer
208
+ p = Path(path)
209
+ if not p.exists():
210
+ from huggingface_hub import snapshot_download
211
+ p = Path(snapshot_download(str(path)))
212
+ fc = json.loads((p / CONFIG_FILE).read_text(encoding="utf-8"))
213
+ ecfg = AutoConfig.from_pretrained(str(p / "encoder"))
214
+ if hasattr(ecfg, "reference_compile"):
215
+ ecfg.reference_compile = False
216
+ try:
217
+ enc = AutoModel.from_config(ecfg, attn_implementation=attn_implementation)
218
+ except Exception:
219
+ enc = AutoModel.from_config(ecfg)
220
+ model = FalconDec(enc, fc)
221
+ wf = p / fc.get("weights", FP16_FILE)
222
+ if not wf.exists():
223
+ wf = p / (INT8_FILE if (p / INT8_FILE).exists() else FP16_FILE)
224
+ sd = load_file(str(wf))
225
+ if wf.name == INT8_FILE:
226
+ sd = dequantize_int8(sd)
227
+ missing, unexpected = model.load_state_dict(sd, strict=False)
228
+ if missing or unexpected:
229
+ print(f"[FalconDec] load warning: missing={list(missing)[:5]} unexpected={list(unexpected)[:5]}")
230
+ tok = AutoTokenizer.from_pretrained(str(p / "tokenizer"))
231
+ dev = torch.device(device) if device is not None else torch.device("cuda" if torch.cuda.is_available() else "cpu")
232
+ model.to(dev)
233
+ if dtype is not None:
234
+ model.to(dtype)
235
+ model.eval()
236
+ return model, tok
237
+
238
+
239
+ # ------------------------------------------------------------------ inference
240
+ def _amp(model):
241
+ if model.device.type == "cuda" and next(model.parameters()).dtype == torch.float32:
242
+ dt = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
243
+ return torch.autocast("cuda", dtype=dt)
244
+ return contextlib.nullcontext()
245
+
246
+
247
+ def _state_text(state):
248
+ if state is None:
249
+ return ""
250
+ return state if isinstance(state, str) else json.dumps(state, ensure_ascii=False)
251
+
252
+
253
+ @torch.no_grad()
254
+ def score_items(model, tok, items, batch_size=32):
255
+ """items: [{"state", "question", "options", "type"?, "option_tokens"?, "seq_len"?}] -> list of prob arrays."""
256
+ fc = model.fcfg
257
+ M = int(fc.get("max_opts_single_pass", 96))
258
+ out = [None] * len(items)
259
+ small = [i for i, it in enumerate(items) if len(it["options"]) <= M]
260
+ for i, it in enumerate(items):
261
+ if len(it["options"]) > M:
262
+ out[i] = _score_large(model, tok, it, batch_size)
263
+ if not small:
264
+ return out
265
+ sp = fc["special"]
266
+ prepared = {}
267
+ for i in small:
268
+ it = items[i]
269
+ if len(it["options"]) < 2:
270
+ raise ValueError("Every question needs at least two options.")
271
+ budget = int(it.get("option_tokens", fc["max_tok_per_opt"]))
272
+ q_ids = tok(it.get("question", ""), add_special_tokens=False, truncation=True, max_length=96)["input_ids"]
273
+ o_ids = tok([str(o) for o in it["options"]], add_special_tokens=False, truncation=True,
274
+ max_length=budget)["input_ids"]
275
+ s = _state_text(it.get("state"))
276
+ s_ids = tok(s, add_special_tokens=False, truncation=True,
277
+ max_length=int(fc["long_max_len"]))["input_ids"] if s else []
278
+ ids, mk = assemble(q_ids, o_ids, s_ids, sp, int(it.get("seq_len", fc["max_len"])), fc["head_max_len"],
279
+ budget, fc["long_max_len"], fc["long_opts_threshold"])
280
+ prepared[i] = (ids, mk, QTYPES[it.get("type", "choice")])
281
+ order = sorted(small, key=lambda i: len(prepared[i][0]))
282
+ model.eval()
283
+ for b0 in range(0, len(order), batch_size):
284
+ idxs = order[b0: b0 + batch_size]
285
+ batch = collate_features([prepared[i] for i in idxs], sp["pad"], model.device)
286
+ with _amp(model):
287
+ logits = model(**batch)
288
+ for j, i in enumerate(idxs):
289
+ n = len(prepared[i][1])
290
+ T = model.temperature[prepared[i][2], n_bucket(n)].float().clamp_min(1e-3)
291
+ out[i] = torch.softmax(logits[j, :n].float() / T, -1).cpu().numpy()
292
+ return out
293
+
294
+
295
+ def _score_large(model, tok, it, batch_size):
296
+ """Tournament for > max_opts_single_pass options: keep the best of each chunk, then one final pass."""
297
+ M = int(model.fcfg.get("max_opts_single_pass", 96))
298
+ opts = list(it["options"])
299
+ cand = list(range(len(opts)))
300
+ while len(cand) > M:
301
+ groups = [cand[i: i + M] for i in range(0, len(cand), M)]
302
+ keep = max(1, M // len(groups))
303
+ ps = score_items(model, tok, [dict(it, options=[opts[c] for c in g]) for g in groups], batch_size)
304
+ cand = [g[t] for g, p in zip(groups, ps) for t in np.argsort(-p)[:keep]]
305
+ p = score_items(model, tok, [dict(it, options=[opts[c] for c in cand])], batch_size)[0]
306
+ full = np.zeros(len(opts), dtype=np.float32)
307
+ full[cand] = p
308
+ return full
309
+
310
+
311
+ def _normalize_question(q):
312
+ qtype = q.get("type", "choice")
313
+ text = q.get("question") or q.get("instructions") or ""
314
+ if qtype == "noul":
315
+ lab = q.get("labels") or {}
316
+ return qtype, text, [True, False], [str(lab.get("true", "Yes")), str(lab.get("false", "No"))]
317
+ crit = q.get("criteria", q.get("options"))
318
+ if qtype == "score":
319
+ opts = [str(c) for c in crit]
320
+ return qtype, text, list(range(len(opts))), opts
321
+ if isinstance(crit, dict):
322
+ return qtype, text, list(crit), [f"{k}: {v}" if v else str(k) for k, v in crit.items()]
323
+ return qtype, text, list(crit), [str(c) for c in crit]
324
+
325
+
326
+ @torch.no_grad()
327
+ def decide(model, tok, state, questions, defer_threshold=None, batch_size=32):
328
+ """Answer typed questions about one state.
329
+
330
+ questions: list of dicts, or a Jev/Laya-style dict {key: question}. Each question:
331
+ {"type": "choice"|"noul"|"score", "question"/"instructions": str,
332
+ "options": [...] or "criteria": {key: description} / [levels], "option_tokens"?: int}
333
+ """
334
+ if isinstance(questions, dict):
335
+ questions = [dict(q, key=k) for k, q in questions.items()]
336
+ items, metas = [], []
337
+ for q in questions:
338
+ qtype, text, keys, opts = _normalize_question(q)
339
+ item = {"state": state, "question": text, "options": opts, "type": qtype}
340
+ for extra in ("option_tokens", "seq_len"):
341
+ if extra in q:
342
+ item[extra] = q[extra]
343
+ items.append(item)
344
+ metas.append((q, qtype, keys, opts))
345
+ probs = score_items(model, tok, items, batch_size)
346
+ thr = model.fcfg.get("defer_threshold", 0.0) if defer_threshold is None else defer_threshold
347
+ results = []
348
+ for (q, qtype, keys, opts), p in zip(metas, probs):
349
+ i = int(np.argmax(p))
350
+ r = {"key": q.get("key"), "type": qtype, "question": items[len(results)]["question"],
351
+ "choice": keys[i], "choice_text": opts[i], "confidence": float(p[i]),
352
+ "probs": {str(k): float(v) for k, v in zip(keys, p)}, "defer": float(p[i]) < thr}
353
+ if qtype == "noul":
354
+ r["p_true"] = float(p[0])
355
+ if qtype == "score":
356
+ r["expected_level"] = float(np.dot(p, np.arange(len(p))))
357
+ results.append(r)
358
+ return {"results": results, "answers": {r["key"]: r for r in results if r["key"] is not None}}
compact-int8/falcondec_report.json ADDED
@@ -0,0 +1,1802 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "version": "1.0.0",
3
+ "notebook_version": "V3.7",
4
+ "mode": "scratch",
5
+ "parent_version": null,
6
+ "lineage": [
7
+ "jhu-clsp/ettin-encoder-150m"
8
+ ],
9
+ "config": {
10
+ "MODE": "scratch",
11
+ "FINETUNE_FROM": "latest",
12
+ "MODEL_NAME": "LightDec_V2_Long",
13
+ "BACKBONE": "jhu-clsp/ettin-encoder-150m",
14
+ "BACKBONE_INIT": "pretrained",
15
+ "PRESET": "long",
16
+ "CUSTOM_DATA_JSONL": "",
17
+ "CUSTOM_WEIGHT": 3.0,
18
+ "INCLUDE_BUILDERS": [],
19
+ "EXCLUDE_BUILDERS": [
20
+ "mind2web"
21
+ ],
22
+ "TEV1_DIR": "./tev1",
23
+ "TEV1_AUTOBUILD": true,
24
+ "TEV1_POLICY_CAP": 13500,
25
+ "TEV1_ROUTING_CAP": 6000,
26
+ "TEV1_RESEARCH_CAP": 1000,
27
+ "TRIAGE_EMAILS": 200,
28
+ "TEV1_RETRIES": 3,
29
+ "TEV1_STEP_TIMEOUT_MIN": 180,
30
+ "TEV1_STALL_MIN": 15,
31
+ "RETRY_FAILED_BUILDERS": true,
32
+ "LONG_TASK_TRAIN": 1500,
33
+ "MAX_LEN": 2048,
34
+ "LONG_MAX_LEN": 2048,
35
+ "LONG_OPTS_THRESHOLD": 24,
36
+ "HEAD_MAX_LEN": 192,
37
+ "MAX_TOK_PER_OPT": 24,
38
+ "MAX_OPTS_SINGLE_PASS": 96,
39
+ "INTERACT_LAYERS": 2,
40
+ "INTERACT_HEADS": 8,
41
+ "BATCH_SIZE": 128,
42
+ "TOKENS_PER_BATCH": 65536,
43
+ "GRAD_ACCUM": 1,
44
+ "LR_ENCODER": 8e-05,
45
+ "LR_HEAD": 0.0006,
46
+ "LLRD": 0.9,
47
+ "WEIGHT_DECAY": 0.01,
48
+ "WARMUP_FRAC": 0.06,
49
+ "EPOCHS": 0,
50
+ "EMA_DECAY": 0.999,
51
+ "GRAD_CLIP": 1.0,
52
+ "FINETUNE_LR_SCALE": 0.5,
53
+ "SPHERICAL_W": 0.5,
54
+ "BRIER_W": 0.0,
55
+ "RPS_W": 1.0,
56
+ "TASK_SAMPLING_ALPHA": 0.5,
57
+ "NOTA_PROB": 0.08,
58
+ "RLCD_EPOCHS": 0.0,
59
+ "RLCD_SAMPLES": 4,
60
+ "RLCD_SIGMA": 0.3,
61
+ "RLCD_CE_ANCHOR": 0.5,
62
+ "DEVICE": "cuda",
63
+ "SEED": 42,
64
+ "NUM_WORKERS": 8,
65
+ "USE_COMPILE": true,
66
+ "ATTN_IMPL": "auto",
67
+ "OUTPUT_ROOT": "./falcondec_runs",
68
+ "EXPORT_INT8": true,
69
+ "SIZE_TARGET_MB": 250,
70
+ "DEFER_THRESHOLD": 0.7,
71
+ "EVAL_BASELINE_PROOF_V2": true,
72
+ "BASELINE_MAX_PER_TASK": 120,
73
+ "RUN_CPU_BENCH": true,
74
+ "PUSH_TO_HUB": false,
75
+ "HUB_REPO": ""
76
+ },
77
+ "epochs": 10,
78
+ "best_epoch": 4,
79
+ "rlcd_kept": false,
80
+ "history": [
81
+ {
82
+ "epoch": 1,
83
+ "train_loss": 0.7058714814313827,
84
+ "train_acc": 0.7342094450703411,
85
+ "val_macro": 0.8188568656508199,
86
+ "val_micro": 0.820702479338843,
87
+ "val_nll": 0.4076740937270036
88
+ },
89
+ {
90
+ "epoch": 2,
91
+ "train_loss": 0.431469622218984,
92
+ "train_acc": 0.8473141073247583,
93
+ "val_macro": 0.8450872188411825,
94
+ "val_micro": 0.8444628099173553,
95
+ "val_nll": 0.3623910443100994
96
+ },
97
+ {
98
+ "epoch": 3,
99
+ "train_loss": 0.34276796075090776,
100
+ "train_acc": 0.8791940521680108,
101
+ "val_macro": 0.8501904675940947,
102
+ "val_micro": 0.8499173553719008,
103
+ "val_nll": 0.39227130460368825
104
+ },
105
+ {
106
+ "epoch": 4,
107
+ "train_loss": 0.2773282952523419,
108
+ "train_acc": 0.9039863196679513,
109
+ "val_macro": 0.8519172452047367,
110
+ "val_micro": 0.8519834710743802,
111
+ "val_nll": 0.4529550437184426
112
+ },
113
+ {
114
+ "epoch": 5,
115
+ "train_loss": 0.22030453415052323,
116
+ "train_acc": 0.925851485588573,
117
+ "val_macro": 0.8488092001804082,
118
+ "val_micro": 0.85,
119
+ "val_nll": 0.5369876145373608
120
+ },
121
+ {
122
+ "epoch": 6,
123
+ "train_loss": 0.16953414903032543,
124
+ "train_acc": 0.9449156076226879,
125
+ "val_macro": 0.8474279008050457,
126
+ "val_micro": 0.8488842975206612,
127
+ "val_nll": 0.6705887193589142
128
+ },
129
+ {
130
+ "epoch": 7,
131
+ "train_loss": 0.126976966611751,
132
+ "train_acc": 0.9602833527848939,
133
+ "val_macro": 0.8453322694829587,
134
+ "val_micro": 0.8476446280991735,
135
+ "val_nll": 0.8430423012207805
136
+ },
137
+ {
138
+ "epoch": 8,
139
+ "train_loss": 0.09470793687766035,
140
+ "train_acc": 0.9714013975720073,
141
+ "val_macro": 0.8451769009269958,
142
+ "val_micro": 0.8467355371900827,
143
+ "val_nll": 1.0336680755654943
144
+ },
145
+ {
146
+ "epoch": 9,
147
+ "train_loss": 0.0710737590306488,
148
+ "train_acc": 0.9794225926173407,
149
+ "val_macro": 0.8435159403610613,
150
+ "val_micro": 0.8458677685950413,
151
+ "val_nll": 1.210751206553627
152
+ },
153
+ {
154
+ "epoch": 10,
155
+ "train_loss": 0.05665995772408678,
156
+ "train_acc": 0.9843832093140154,
157
+ "val_macro": 0.844272704258198,
158
+ "val_micro": 0.8474380165289256,
159
+ "val_nll": 1.41872131482475
160
+ }
161
+ ],
162
+ "train_minutes": 291.67019440333047,
163
+ "temperatures": [
164
+ [
165
+ 2.37591814994812,
166
+ 2.0426506996154785,
167
+ 1.6455482244491577,
168
+ 1.8417285680770874
169
+ ],
170
+ [
171
+ 2.4885058403015137,
172
+ 2.4885058403015137,
173
+ 2.4885058403015137,
174
+ 2.4885058403015137
175
+ ],
176
+ [
177
+ 2.2925639152526855,
178
+ 2.2925639152526855,
179
+ 2.2925639152526855,
180
+ 2.2925639152526855
181
+ ]
182
+ ],
183
+ "data_counts": {
184
+ "train": 1023814,
185
+ "validation": 24200,
186
+ "test": 31990
187
+ },
188
+ "skipped_builders": [],
189
+ "test_overall": {
190
+ "n": 31990.0,
191
+ "micro_acc": 0.784057517974367,
192
+ "macro_acc": 0.7887006006875437,
193
+ "nll": 0.5630830966790886,
194
+ "brier": 0.2954194093899411,
195
+ "ece": 0.02856345933167103,
196
+ "aurc": 0.06778106632931055,
197
+ "score_mae": 0.5060575584353655,
198
+ "coverage@0.7": 0.6797436698968428,
199
+ "acc_on_covered@0.7": 0.9059094044607956
200
+ },
201
+ "test_per_domain": {
202
+ "agentic": 0.821117397177404,
203
+ "classification": 0.6115,
204
+ "code": 0.9108344674425007,
205
+ "guardrails": 0.7984073149310548,
206
+ "intents": 0.8505,
207
+ "long_context": 1.0,
208
+ "policy": 0.8061818181818182,
209
+ "reasoning": 0.7180877154615182,
210
+ "support": 0.9815181518151815,
211
+ "tev1_benchmark": 0.6541666666666667,
212
+ "workflows": 0.773
213
+ },
214
+ "test_per_task": [
215
+ {
216
+ "task": "ag_news/topic",
217
+ "n": 500,
218
+ "acc": 0.9,
219
+ "chance": 0.25,
220
+ "nll": 0.2674782864307199,
221
+ "brier": 0.14296993909824685,
222
+ "heldout": false,
223
+ "domain": "classification",
224
+ "ece": 0.05165432274341582,
225
+ "proof_v2 (card)": NaN,
226
+ "\u0394 vs proof_v2": NaN
227
+ },
228
+ {
229
+ "task": "agentharm/refuse",
230
+ "n": 416,
231
+ "acc": 0.6995192307692307,
232
+ "chance": 0.5,
233
+ "nll": 0.6008703433729422,
234
+ "brier": 0.4076762512706606,
235
+ "heldout": true,
236
+ "domain": "guardrails",
237
+ "ece": 0.08916855761064935,
238
+ "proof_v2 (card)": NaN,
239
+ "\u0394 vs proof_v2": NaN
240
+ },
241
+ {
242
+ "task": "agenttrek/finish_now",
243
+ "n": 151,
244
+ "acc": 0.7880794701986755,
245
+ "chance": 0.5,
246
+ "nll": 0.445978463758558,
247
+ "brier": 0.30066953632150406,
248
+ "heldout": false,
249
+ "domain": "agentic",
250
+ "ece": 0.07512128787324918,
251
+ "proof_v2 (card)": NaN,
252
+ "\u0394 vs proof_v2": NaN
253
+ },
254
+ {
255
+ "task": "agenttrek/next_action_type",
256
+ "n": 487,
257
+ "acc": 0.864476386036961,
258
+ "chance": 0.19054952576513148,
259
+ "nll": 0.3929848723152591,
260
+ "brier": 0.2184262054212701,
261
+ "heldout": false,
262
+ "domain": "agentic",
263
+ "ece": 0.09202881025827395,
264
+ "proof_v2 (card)": NaN,
265
+ "\u0394 vs proof_v2": NaN
266
+ },
267
+ {
268
+ "task": "anli/nli",
269
+ "n": 498,
270
+ "acc": 0.4859437751004016,
271
+ "chance": 0.33333333333333326,
272
+ "nll": 1.069221544292677,
273
+ "brier": 0.6544603492386665,
274
+ "heldout": false,
275
+ "domain": "reasoning",
276
+ "ece": 0.14589204983299517,
277
+ "proof_v2 (card)": NaN,
278
+ "\u0394 vs proof_v2": NaN
279
+ },
280
+ {
281
+ "task": "aqua_rat/math",
282
+ "n": 247,
283
+ "acc": 0.340080971659919,
284
+ "chance": 0.2,
285
+ "nll": 1.54525283087603,
286
+ "brier": 0.7739737508966391,
287
+ "heldout": false,
288
+ "domain": "reasoning",
289
+ "ece": 0.080350092668765,
290
+ "proof_v2 (card)": NaN,
291
+ "\u0394 vs proof_v2": NaN
292
+ },
293
+ {
294
+ "task": "arc_challenge/mcq",
295
+ "n": 500,
296
+ "acc": 0.49,
297
+ "chance": 0.25006666666666666,
298
+ "nll": 1.3021090898208785,
299
+ "brier": 0.6752390805252713,
300
+ "heldout": true,
301
+ "domain": "reasoning",
302
+ "ece": 0.1680490539073944,
303
+ "proof_v2 (card)": 0.308,
304
+ "\u0394 vs proof_v2": 0.182
305
+ },
306
+ {
307
+ "task": "arc_easy/mcq",
308
+ "n": 500,
309
+ "acc": 0.612,
310
+ "chance": 0.24980000000000002,
311
+ "nll": 0.9830388274090365,
312
+ "brier": 0.5264415342806388,
313
+ "heldout": true,
314
+ "domain": "reasoning",
315
+ "ece": 0.1215095258355141,
316
+ "proof_v2 (card)": 0.425,
317
+ "\u0394 vs proof_v2": 0.187
318
+ },
319
+ {
320
+ "task": "banking77/intent",
321
+ "n": 500,
322
+ "acc": 0.922,
323
+ "chance": 0.3214666666666666,
324
+ "nll": 0.23128961177584006,
325
+ "brier": 0.11756106243365176,
326
+ "heldout": true,
327
+ "domain": "intents",
328
+ "ece": 0.037492330908775316,
329
+ "proof_v2 (card)": 0.883,
330
+ "\u0394 vs proof_v2": 0.039000000000000035
331
+ },
332
+ {
333
+ "task": "banking77/intent_77",
334
+ "n": 300,
335
+ "acc": 0.58,
336
+ "chance": 0.012987012987012986,
337
+ "nll": 2.159930980304489,
338
+ "brier": 0.6431857202493357,
339
+ "heldout": true,
340
+ "domain": "intents",
341
+ "ece": 0.24079640249411272,
342
+ "proof_v2 (card)": NaN,
343
+ "\u0394 vs proof_v2": NaN
344
+ },
345
+ {
346
+ "task": "bigclonebench/clone",
347
+ "n": 500,
348
+ "acc": 0.962,
349
+ "chance": 0.5,
350
+ "nll": 0.1312969167594565,
351
+ "brier": 0.06444151702850223,
352
+ "heldout": false,
353
+ "domain": "code",
354
+ "ece": 0.03135248208045959,
355
+ "proof_v2 (card)": 0.383,
356
+ "\u0394 vs proof_v2": 0.579
357
+ },
358
+ {
359
+ "task": "bitext/category",
360
+ "n": 500,
361
+ "acc": 1.0,
362
+ "chance": 0.16128124098124094,
363
+ "nll": 0.0013673773490934594,
364
+ "brier": 0.0003828941821233322,
365
+ "heldout": false,
366
+ "domain": "support",
367
+ "ece": 0.001247308015823329,
368
+ "proof_v2 (card)": NaN,
369
+ "\u0394 vs proof_v2": NaN
370
+ },
371
+ {
372
+ "task": "bitext/route",
373
+ "n": 500,
374
+ "acc": 1.0,
375
+ "chance": 0.19016313131313134,
376
+ "nll": 0.00031508885690954,
377
+ "brier": 4.903548307607296e-06,
378
+ "heldout": false,
379
+ "domain": "support",
380
+ "ece": 0.0003137041330337764,
381
+ "proof_v2 (card)": 0.958,
382
+ "\u0394 vs proof_v2": 0.04200000000000004
383
+ },
384
+ {
385
+ "task": "boolq/yes_no",
386
+ "n": 500,
387
+ "acc": 0.822,
388
+ "chance": 0.5,
389
+ "nll": 0.4842173567947466,
390
+ "brier": 0.2840138305161832,
391
+ "heldout": false,
392
+ "domain": "reasoning",
393
+ "ece": 0.0795972980260849,
394
+ "proof_v2 (card)": 0.717,
395
+ "\u0394 vs proof_v2": 0.10499999999999998
396
+ },
397
+ {
398
+ "task": "civil_comments/toxic",
399
+ "n": 500,
400
+ "acc": 0.926,
401
+ "chance": 0.5,
402
+ "nll": 0.2234244500696659,
403
+ "brier": 0.1209681824839501,
404
+ "heldout": false,
405
+ "domain": "guardrails",
406
+ "ece": 0.05123518884181974,
407
+ "proof_v2 (card)": NaN,
408
+ "\u0394 vs proof_v2": NaN
409
+ },
410
+ {
411
+ "task": "clinc150/intent",
412
+ "n": 500,
413
+ "acc": 0.97,
414
+ "chance": 0.12404474984829318,
415
+ "nll": 0.09906445816022005,
416
+ "brier": 0.04996487012600131,
417
+ "heldout": false,
418
+ "domain": "intents",
419
+ "ece": 0.014944501101970714,
420
+ "proof_v2 (card)": 0.85,
421
+ "\u0394 vs proof_v2": 0.12
422
+ },
423
+ {
424
+ "task": "codexglue/code_to_doc",
425
+ "n": 504,
426
+ "acc": 0.9900793650793651,
427
+ "chance": 0.26233465608465617,
428
+ "nll": 0.032474144255576765,
429
+ "brier": 0.016158895201785088,
430
+ "heldout": false,
431
+ "domain": "code",
432
+ "ece": 0.009654374349685009,
433
+ "proof_v2 (card)": 0.969,
434
+ "\u0394 vs proof_v2": 0.021079365079365142
435
+ },
436
+ {
437
+ "task": "codexglue/doc_to_code",
438
+ "n": 504,
439
+ "acc": 0.9761904761904762,
440
+ "chance": 0.2786044973544973,
441
+ "nll": 0.05970113779317181,
442
+ "brier": 0.03072688989396561,
443
+ "heldout": false,
444
+ "domain": "code",
445
+ "ece": 0.012547755761752008,
446
+ "proof_v2 (card)": 0.961,
447
+ "\u0394 vs proof_v2": 0.015190476190476199
448
+ },
449
+ {
450
+ "task": "codexglue/func_name",
451
+ "n": 467,
452
+ "acc": 0.9550321199143469,
453
+ "chance": 0.25499643112062814,
454
+ "nll": 0.14961532151979784,
455
+ "brier": 0.07053876912275295,
456
+ "heldout": false,
457
+ "domain": "code",
458
+ "ece": 0.03948740715388316,
459
+ "proof_v2 (card)": 0.901,
460
+ "\u0394 vs proof_v2": 0.054032119914346866
461
+ },
462
+ {
463
+ "task": "codexglue/lang_id",
464
+ "n": 504,
465
+ "acc": 1.0,
466
+ "chance": 0.2354166666666667,
467
+ "nll": 0.0014352430686462487,
468
+ "brier": 1.2864369183890018e-05,
469
+ "heldout": false,
470
+ "domain": "code",
471
+ "ece": 0.0014309174721203188,
472
+ "proof_v2 (card)": 0.997,
473
+ "\u0394 vs proof_v2": 0.0030000000000000027
474
+ },
475
+ {
476
+ "task": "commonsense_qa/mcq",
477
+ "n": 493,
478
+ "acc": 0.6430020283975659,
479
+ "chance": 0.19999999999999998,
480
+ "nll": 0.9523430099929484,
481
+ "brier": 0.49435087368161623,
482
+ "heldout": false,
483
+ "domain": "reasoning",
484
+ "ece": 0.14308575696442233,
485
+ "proof_v2 (card)": 0.442,
486
+ "\u0394 vs proof_v2": 0.20100202839756592
487
+ },
488
+ {
489
+ "task": "counsel/critique_quality",
490
+ "n": 201,
491
+ "acc": 0.6268656716417911,
492
+ "chance": 0.3333333333333333,
493
+ "nll": 1.5466346353110312,
494
+ "brier": 0.6427543071457474,
495
+ "heldout": false,
496
+ "domain": "agentic",
497
+ "ece": 0.27062572442477023,
498
+ "proof_v2 (card)": NaN,
499
+ "\u0394 vs proof_v2": NaN
500
+ },
501
+ {
502
+ "task": "counsel/step_has_error",
503
+ "n": 201,
504
+ "acc": 0.8308457711442786,
505
+ "chance": 0.5,
506
+ "nll": 0.8236887436017318,
507
+ "brier": 0.3201374357369142,
508
+ "heldout": false,
509
+ "domain": "agentic",
510
+ "ece": 0.14827988011326954,
511
+ "proof_v2 (card)": NaN,
512
+ "\u0394 vs proof_v2": NaN
513
+ },
514
+ {
515
+ "task": "devign/vulnerability",
516
+ "n": 500,
517
+ "acc": 0.632,
518
+ "chance": 0.5,
519
+ "nll": 0.6074751503933221,
520
+ "brier": 0.42706214521965163,
521
+ "heldout": false,
522
+ "domain": "code",
523
+ "ece": 0.07018342161178594,
524
+ "proof_v2 (card)": 0.542,
525
+ "\u0394 vs proof_v2": 0.08999999999999997
526
+ },
527
+ {
528
+ "task": "emotion/6way",
529
+ "n": 500,
530
+ "acc": 0.472,
531
+ "chance": 0.16666666666666663,
532
+ "nll": 1.561316128242761,
533
+ "brier": 0.7356062387487318,
534
+ "heldout": true,
535
+ "domain": "classification",
536
+ "ece": 0.21054179659485817,
537
+ "proof_v2 (card)": NaN,
538
+ "\u0394 vs proof_v2": NaN
539
+ },
540
+ {
541
+ "task": "gsm8k/math",
542
+ "n": 500,
543
+ "acc": 0.7,
544
+ "chance": 0.25,
545
+ "nll": 0.6923719964642078,
546
+ "brier": 0.4014884319624292,
547
+ "heldout": false,
548
+ "domain": "reasoning",
549
+ "ece": 0.08541783547401426,
550
+ "proof_v2 (card)": 0.275,
551
+ "\u0394 vs proof_v2": 0.42499999999999993
552
+ },
553
+ {
554
+ "task": "hellaswag/continuation",
555
+ "n": 500,
556
+ "acc": 0.576,
557
+ "chance": 0.25,
558
+ "nll": 0.968790470642969,
559
+ "brier": 0.5288133586177948,
560
+ "heldout": false,
561
+ "domain": "reasoning",
562
+ "ece": 0.07543198531866073,
563
+ "proof_v2 (card)": NaN,
564
+ "\u0394 vs proof_v2": NaN
565
+ },
566
+ {
567
+ "task": "hotpotqa/comparison_yes_no",
568
+ "n": 26,
569
+ "acc": 0.9230769230769231,
570
+ "chance": 0.5,
571
+ "nll": 0.3400710394176153,
572
+ "brier": 0.15484679762361436,
573
+ "heldout": false,
574
+ "domain": "agentic",
575
+ "ece": 0.08232885369887719,
576
+ "proof_v2 (card)": NaN,
577
+ "\u0394 vs proof_v2": NaN
578
+ },
579
+ {
580
+ "task": "hotpotqa/retrieve",
581
+ "n": 497,
582
+ "acc": 0.8933601609657947,
583
+ "chance": 0.16741720162243304,
584
+ "nll": 0.3457089332123268,
585
+ "brier": 0.16216258689991922,
586
+ "heldout": false,
587
+ "domain": "agentic",
588
+ "ece": 0.04268557528854612,
589
+ "proof_v2 (card)": NaN,
590
+ "\u0394 vs proof_v2": NaN
591
+ },
592
+ {
593
+ "task": "humaneval/completion",
594
+ "n": 119,
595
+ "acc": 0.8487394957983193,
596
+ "chance": 0.39355742296918783,
597
+ "nll": 0.46361927195851294,
598
+ "brier": 0.2730045795813753,
599
+ "heldout": true,
600
+ "domain": "code",
601
+ "ece": 0.13989568708323633,
602
+ "proof_v2 (card)": 0.575,
603
+ "\u0394 vs proof_v2": 0.27373949579831935
604
+ },
605
+ {
606
+ "task": "jailbreak/detect",
607
+ "n": 262,
608
+ "acc": 0.9732824427480916,
609
+ "chance": 0.5,
610
+ "nll": 0.08469655378148465,
611
+ "brier": 0.044003931208438284,
612
+ "heldout": false,
613
+ "domain": "guardrails",
614
+ "ece": 0.030764950368240628,
615
+ "proof_v2 (card)": NaN,
616
+ "\u0394 vs proof_v2": NaN
617
+ },
618
+ {
619
+ "task": "long/contract_clause",
620
+ "n": 150,
621
+ "acc": 1.0,
622
+ "chance": 0.19999999999999996,
623
+ "nll": 0.0002789537957808837,
624
+ "brier": 4.274148862877054e-07,
625
+ "heldout": false,
626
+ "domain": "long_context",
627
+ "ece": 0.00027882695198055973,
628
+ "proof_v2 (card)": NaN,
629
+ "\u0394 vs proof_v2": NaN
630
+ },
631
+ {
632
+ "task": "long/email_thread",
633
+ "n": 150,
634
+ "acc": 1.0,
635
+ "chance": 0.25,
636
+ "nll": 0.0003266528345905802,
637
+ "brier": 2.717841913247229e-06,
638
+ "heldout": false,
639
+ "domain": "long_context",
640
+ "ece": 0.0003259424368540209,
641
+ "proof_v2 (card)": NaN,
642
+ "\u0394 vs proof_v2": NaN
643
+ },
644
+ {
645
+ "task": "long/service_log",
646
+ "n": 150,
647
+ "acc": 1.0,
648
+ "chance": 0.19999999999999996,
649
+ "nll": 0.0010035930563268873,
650
+ "brier": 2.0987077842670875e-06,
651
+ "heldout": false,
652
+ "domain": "long_context",
653
+ "ece": 0.0010028688112894146,
654
+ "proof_v2 (card)": NaN,
655
+ "\u0394 vs proof_v2": NaN
656
+ },
657
+ {
658
+ "task": "massive_en/intent",
659
+ "n": 500,
660
+ "acc": 0.93,
661
+ "chance": 0.12958727789085603,
662
+ "nll": 0.20636169074449573,
663
+ "brier": 0.10375277943487987,
664
+ "heldout": false,
665
+ "domain": "intents",
666
+ "ece": 0.0284088468849659,
667
+ "proof_v2 (card)": NaN,
668
+ "\u0394 vs proof_v2": NaN
669
+ },
670
+ {
671
+ "task": "mbpp/bugspot",
672
+ "n": 256,
673
+ "acc": 0.85546875,
674
+ "chance": 0.40559895833333326,
675
+ "nll": 0.3667814989712497,
676
+ "brier": 0.21843909652329055,
677
+ "heldout": false,
678
+ "domain": "code",
679
+ "ece": 0.037241218378767364,
680
+ "proof_v2 (card)": 0.475,
681
+ "\u0394 vs proof_v2": 0.38046875
682
+ },
683
+ {
684
+ "task": "mbpp/solution",
685
+ "n": 500,
686
+ "acc": 0.978,
687
+ "chance": 0.25,
688
+ "nll": 0.05863042902501184,
689
+ "brier": 0.02981698831494921,
690
+ "heldout": false,
691
+ "domain": "code",
692
+ "ece": 0.01960155099630357,
693
+ "proof_v2 (card)": 0.992,
694
+ "\u0394 vs proof_v2": -0.014000000000000012
695
+ },
696
+ {
697
+ "task": "mmlu/mcq",
698
+ "n": 500,
699
+ "acc": 0.39,
700
+ "chance": 0.25,
701
+ "nll": 1.4543649319559335,
702
+ "brier": 0.7702575193705765,
703
+ "heldout": true,
704
+ "domain": "reasoning",
705
+ "ece": 0.15507307499647138,
706
+ "proof_v2 (card)": NaN,
707
+ "\u0394 vs proof_v2": NaN
708
+ },
709
+ {
710
+ "task": "mnli/claim",
711
+ "n": 500,
712
+ "acc": 0.86,
713
+ "chance": 0.33333333333333326,
714
+ "nll": 0.4006310784481466,
715
+ "brier": 0.2131232579488494,
716
+ "heldout": false,
717
+ "domain": "reasoning",
718
+ "ece": 0.081218329668045,
719
+ "proof_v2 (card)": 0.492,
720
+ "\u0394 vs proof_v2": 0.368
721
+ },
722
+ {
723
+ "task": "openbookqa/mcq",
724
+ "n": 500,
725
+ "acc": 0.572,
726
+ "chance": 0.25,
727
+ "nll": 1.1920220527790952,
728
+ "brier": 0.6155284183368397,
729
+ "heldout": false,
730
+ "domain": "reasoning",
731
+ "ece": 0.21775564682483672,
732
+ "proof_v2 (card)": 0.292,
733
+ "\u0394 vs proof_v2": 0.27999999999999997
734
+ },
735
+ {
736
+ "task": "policy/access_control_transfer",
737
+ "n": 500,
738
+ "acc": 1.0,
739
+ "chance": 0.33333333333333326,
740
+ "nll": 0.002006467206090747,
741
+ "brier": 2.4728843418299015e-05,
742
+ "heldout": false,
743
+ "domain": "policy",
744
+ "ece": 0.001999327063560541,
745
+ "proof_v2 (card)": NaN,
746
+ "\u0394 vs proof_v2": NaN
747
+ },
748
+ {
749
+ "task": "policy/count_threshold_transfer",
750
+ "n": 500,
751
+ "acc": 0.88,
752
+ "chance": 0.10537052392052391,
753
+ "nll": 0.319376940273738,
754
+ "brier": 0.1809248790341789,
755
+ "heldout": false,
756
+ "domain": "policy",
757
+ "ece": 0.04758520478010175,
758
+ "proof_v2 (card)": NaN,
759
+ "\u0394 vs proof_v2": NaN
760
+ },
761
+ {
762
+ "task": "policy/free_shipping_transfer",
763
+ "n": 500,
764
+ "acc": 0.938,
765
+ "chance": 0.5,
766
+ "nll": 0.1287872482436942,
767
+ "brier": 0.08322479293591872,
768
+ "heldout": false,
769
+ "domain": "policy",
770
+ "ece": 0.041308605194091734,
771
+ "proof_v2 (card)": NaN,
772
+ "\u0394 vs proof_v2": NaN
773
+ },
774
+ {
775
+ "task": "policy/invoice_overdue_transfer",
776
+ "n": 500,
777
+ "acc": 0.824,
778
+ "chance": 0.5,
779
+ "nll": 0.3965394942490384,
780
+ "brier": 0.2504247577332596,
781
+ "heldout": false,
782
+ "domain": "policy",
783
+ "ece": 0.01985593855381013,
784
+ "proof_v2 (card)": NaN,
785
+ "\u0394 vs proof_v2": NaN
786
+ },
787
+ {
788
+ "task": "policy/invoice_total_transfer",
789
+ "n": 500,
790
+ "acc": 0.474,
791
+ "chance": 0.5,
792
+ "nll": 0.710681935429573,
793
+ "brier": 0.5164436840846652,
794
+ "heldout": false,
795
+ "domain": "policy",
796
+ "ece": 0.08196303224563599,
797
+ "proof_v2 (card)": NaN,
798
+ "\u0394 vs proof_v2": NaN
799
+ },
800
+ {
801
+ "task": "policy/refund_approval_transfer",
802
+ "n": 500,
803
+ "acc": 0.962,
804
+ "chance": 0.33333333333333326,
805
+ "nll": 0.10006511305032836,
806
+ "brier": 0.057969574647252664,
807
+ "heldout": false,
808
+ "domain": "policy",
809
+ "ece": 0.018834196925163266,
810
+ "proof_v2 (card)": NaN,
811
+ "\u0394 vs proof_v2": NaN
812
+ },
813
+ {
814
+ "task": "policy/return_window_transfer",
815
+ "n": 500,
816
+ "acc": 1.0,
817
+ "chance": 0.33333333333333326,
818
+ "nll": 0.02710788446944207,
819
+ "brier": 0.004185976644178241,
820
+ "heldout": false,
821
+ "domain": "policy",
822
+ "ece": 0.02567652928829193,
823
+ "proof_v2 (card)": NaN,
824
+ "\u0394 vs proof_v2": NaN
825
+ },
826
+ {
827
+ "task": "policy/sla_urgency_transfer",
828
+ "n": 500,
829
+ "acc": 0.606,
830
+ "chance": 0.25,
831
+ "nll": 0.8170870101451874,
832
+ "brier": 0.5123369210022903,
833
+ "heldout": false,
834
+ "domain": "policy",
835
+ "ece": 0.13068743020296097,
836
+ "proof_v2 (card)": NaN,
837
+ "\u0394 vs proof_v2": NaN
838
+ },
839
+ {
840
+ "task": "policy/table_compare_transfer",
841
+ "n": 500,
842
+ "acc": 0.928,
843
+ "chance": 0.5,
844
+ "nll": 0.2105378701629961,
845
+ "brier": 0.11898007306418568,
846
+ "heldout": false,
847
+ "domain": "policy",
848
+ "ece": 0.04929606854915622,
849
+ "proof_v2 (card)": NaN,
850
+ "\u0394 vs proof_v2": NaN
851
+ },
852
+ {
853
+ "task": "policy/table_count_transfer",
854
+ "n": 500,
855
+ "acc": 0.324,
856
+ "chance": 0.12114228549228548,
857
+ "nll": 3.24442003638536,
858
+ "brier": 1.1488026091975905,
859
+ "heldout": false,
860
+ "domain": "policy",
861
+ "ece": 0.5407210965156555,
862
+ "proof_v2 (card)": NaN,
863
+ "\u0394 vs proof_v2": NaN
864
+ },
865
+ {
866
+ "task": "policy/table_extreme_transfer",
867
+ "n": 500,
868
+ "acc": 0.932,
869
+ "chance": 0.12210815295815294,
870
+ "nll": 0.268062693382595,
871
+ "brier": 0.11322196827103552,
872
+ "heldout": false,
873
+ "domain": "policy",
874
+ "ece": 0.05259446609020234,
875
+ "proof_v2 (card)": NaN,
876
+ "\u0394 vs proof_v2": NaN
877
+ },
878
+ {
879
+ "task": "prompt_injections/detect",
880
+ "n": 116,
881
+ "acc": 0.5948275862068966,
882
+ "chance": 0.5,
883
+ "nll": 1.1838965486195179,
884
+ "brier": 0.6826645118148511,
885
+ "heldout": true,
886
+ "domain": "guardrails",
887
+ "ece": 0.33272201953263114,
888
+ "proof_v2 (card)": NaN,
889
+ "\u0394 vs proof_v2": NaN
890
+ },
891
+ {
892
+ "task": "qasc/mcq",
893
+ "n": 500,
894
+ "acc": 0.986,
895
+ "chance": 0.125,
896
+ "nll": 0.028506858559036514,
897
+ "brier": 0.017274368003708154,
898
+ "heldout": false,
899
+ "domain": "reasoning",
900
+ "ece": 0.007774531126022318,
901
+ "proof_v2 (card)": NaN,
902
+ "\u0394 vs proof_v2": NaN
903
+ },
904
+ {
905
+ "task": "sciq/mcq",
906
+ "n": 498,
907
+ "acc": 0.9538152610441767,
908
+ "chance": 0.25,
909
+ "nll": 0.1506263591345834,
910
+ "brier": 0.06985463059035983,
911
+ "heldout": false,
912
+ "domain": "reasoning",
913
+ "ece": 0.020590302994452258,
914
+ "proof_v2 (card)": 0.692,
915
+ "\u0394 vs proof_v2": 0.26181526104417674
916
+ },
917
+ {
918
+ "task": "scitail/support",
919
+ "n": 500,
920
+ "acc": 0.962,
921
+ "chance": 0.5,
922
+ "nll": 0.11612921302905306,
923
+ "brier": 0.05990850379879065,
924
+ "heldout": false,
925
+ "domain": "reasoning",
926
+ "ece": 0.03015338957309728,
927
+ "proof_v2 (card)": NaN,
928
+ "\u0394 vs proof_v2": NaN
929
+ },
930
+ {
931
+ "task": "snli/contradicts",
932
+ "n": 500,
933
+ "acc": 0.99,
934
+ "chance": 0.33333333333333326,
935
+ "nll": 0.042843449617153966,
936
+ "brier": 0.017027529190616314,
937
+ "heldout": false,
938
+ "domain": "reasoning",
939
+ "ece": 0.018511961221695038,
940
+ "proof_v2 (card)": 0.892,
941
+ "\u0394 vs proof_v2": 0.09799999999999998
942
+ },
943
+ {
944
+ "task": "snli/must_be_true",
945
+ "n": 500,
946
+ "acc": 0.986,
947
+ "chance": 0.33333333333333326,
948
+ "nll": 0.06753946821717545,
949
+ "brier": 0.02560326154020915,
950
+ "heldout": false,
951
+ "domain": "reasoning",
952
+ "ece": 0.043202445626258815,
953
+ "proof_v2 (card)": 0.908,
954
+ "\u0394 vs proof_v2": 0.07799999999999996
955
+ },
956
+ {
957
+ "task": "snli/nli",
958
+ "n": 988,
959
+ "acc": 0.8947368421052632,
960
+ "chance": 0.33333333333333326,
961
+ "nll": 0.3554420507991845,
962
+ "brier": 0.17963376213248355,
963
+ "heldout": false,
964
+ "domain": "reasoning",
965
+ "ece": 0.10791712172842222,
966
+ "proof_v2 (card)": NaN,
967
+ "\u0394 vs proof_v2": NaN
968
+ },
969
+ {
970
+ "task": "sst5/score",
971
+ "n": 500,
972
+ "acc": 0.406,
973
+ "chance": 0.2,
974
+ "nll": 1.3002369542717933,
975
+ "brier": 0.6888741553408958,
976
+ "heldout": true,
977
+ "domain": "classification",
978
+ "ece": 0.068851743131876,
979
+ "proof_v2 (card)": NaN,
980
+ "\u0394 vs proof_v2": NaN
981
+ },
982
+ {
983
+ "task": "tev1_test/ag_news",
984
+ "n": 150,
985
+ "acc": 0.92,
986
+ "chance": 0.25,
987
+ "nll": 0.2815313008365532,
988
+ "brier": 0.13456139206846243,
989
+ "heldout": true,
990
+ "domain": "tev1_benchmark",
991
+ "ece": 0.09770219445228576,
992
+ "proof_v2 (card)": NaN,
993
+ "\u0394 vs proof_v2": NaN
994
+ },
995
+ {
996
+ "task": "tev1_test/banking77",
997
+ "n": 200,
998
+ "acc": 0.72,
999
+ "chance": 0.17791666666666664,
1000
+ "nll": 0.777236735031438,
1001
+ "brier": 0.3902663054271286,
1002
+ "heldout": true,
1003
+ "domain": "tev1_benchmark",
1004
+ "ece": 0.13721376925706863,
1005
+ "proof_v2 (card)": NaN,
1006
+ "\u0394 vs proof_v2": NaN
1007
+ },
1008
+ {
1009
+ "task": "tev1_test/boolq",
1010
+ "n": 200,
1011
+ "acc": 0.835,
1012
+ "chance": 0.5,
1013
+ "nll": 0.40902234488283284,
1014
+ "brier": 0.24753055362669119,
1015
+ "heldout": true,
1016
+ "domain": "tev1_benchmark",
1017
+ "ece": 0.04955918818712238,
1018
+ "proof_v2 (card)": NaN,
1019
+ "\u0394 vs proof_v2": NaN
1020
+ },
1021
+ {
1022
+ "task": "tev1_test/mnli",
1023
+ "n": 300,
1024
+ "acc": 0.7233333333333334,
1025
+ "chance": 0.3333333333333333,
1026
+ "nll": 0.6736717029288412,
1027
+ "brier": 0.39782655881245704,
1028
+ "heldout": true,
1029
+ "domain": "tev1_benchmark",
1030
+ "ece": 0.07533363938331603,
1031
+ "proof_v2 (card)": NaN,
1032
+ "\u0394 vs proof_v2": NaN
1033
+ },
1034
+ {
1035
+ "task": "tev1_test/policy",
1036
+ "n": 1200,
1037
+ "acc": 0.5291666666666667,
1038
+ "chance": 0.3333333333333333,
1039
+ "nll": 0.8761954319352905,
1040
+ "brier": 0.5462664224307302,
1041
+ "heldout": true,
1042
+ "domain": "tev1_benchmark",
1043
+ "ece": 0.05868758827447889,
1044
+ "proof_v2 (card)": NaN,
1045
+ "\u0394 vs proof_v2": NaN
1046
+ },
1047
+ {
1048
+ "task": "tev1_test/routing",
1049
+ "n": 600,
1050
+ "acc": 0.4583333333333333,
1051
+ "chance": 0.2,
1052
+ "nll": 1.120521725914441,
1053
+ "brier": 0.5789269511485964,
1054
+ "heldout": true,
1055
+ "domain": "tev1_benchmark",
1056
+ "ece": 0.037736288358767814,
1057
+ "proof_v2 (card)": NaN,
1058
+ "\u0394 vs proof_v2": NaN
1059
+ },
1060
+ {
1061
+ "task": "tev1_test/sst5",
1062
+ "n": 150,
1063
+ "acc": 0.3933333333333333,
1064
+ "chance": 0.19999999999999996,
1065
+ "nll": 1.294863009850184,
1066
+ "brier": 0.6810328679494898,
1067
+ "heldout": true,
1068
+ "domain": "tev1_benchmark",
1069
+ "ece": 0.1081918982664744,
1070
+ "proof_v2 (card)": NaN,
1071
+ "\u0394 vs proof_v2": NaN
1072
+ },
1073
+ {
1074
+ "task": "triage/support_email",
1075
+ "n": 505,
1076
+ "acc": 0.9445544554455445,
1077
+ "chance": 0.32333333333333325,
1078
+ "nll": 0.15705669974757994,
1079
+ "brier": 0.07848840114983771,
1080
+ "heldout": false,
1081
+ "domain": "support",
1082
+ "ece": 0.045869263741049465,
1083
+ "proof_v2 (card)": NaN,
1084
+ "\u0394 vs proof_v2": NaN
1085
+ },
1086
+ {
1087
+ "task": "typed_decisions/agent_trace_observability",
1088
+ "n": 500,
1089
+ "acc": 0.734,
1090
+ "chance": 0.3,
1091
+ "nll": 0.8089503426551818,
1092
+ "brier": 0.45125534421065666,
1093
+ "heldout": false,
1094
+ "domain": "workflows",
1095
+ "ece": 0.22166652178764346,
1096
+ "proof_v2 (card)": NaN,
1097
+ "\u0394 vs proof_v2": NaN
1098
+ },
1099
+ {
1100
+ "task": "typed_decisions/customer_service",
1101
+ "n": 500,
1102
+ "acc": 0.764,
1103
+ "chance": 0.28,
1104
+ "nll": 0.7428050636351109,
1105
+ "brier": 0.40177602517263195,
1106
+ "heldout": false,
1107
+ "domain": "workflows",
1108
+ "ece": 0.20972245779633528,
1109
+ "proof_v2 (card)": NaN,
1110
+ "\u0394 vs proof_v2": NaN
1111
+ },
1112
+ {
1113
+ "task": "typed_decisions/invoice_processing",
1114
+ "n": 500,
1115
+ "acc": 0.826,
1116
+ "chance": 0.35,
1117
+ "nll": 0.5732684296518564,
1118
+ "brier": 0.2997603036136731,
1119
+ "heldout": false,
1120
+ "domain": "workflows",
1121
+ "ece": 0.19463082921504973,
1122
+ "proof_v2 (card)": NaN,
1123
+ "\u0394 vs proof_v2": NaN
1124
+ },
1125
+ {
1126
+ "task": "typed_decisions/security_incidents",
1127
+ "n": 500,
1128
+ "acc": 0.768,
1129
+ "chance": 0.34,
1130
+ "nll": 0.7502468670606613,
1131
+ "brier": 0.4238190137148829,
1132
+ "heldout": false,
1133
+ "domain": "workflows",
1134
+ "ece": 0.23508695220947262,
1135
+ "proof_v2 (card)": NaN,
1136
+ "\u0394 vs proof_v2": NaN
1137
+ },
1138
+ {
1139
+ "task": "winogrande/blank",
1140
+ "n": 500,
1141
+ "acc": 0.662,
1142
+ "chance": 0.5,
1143
+ "nll": 0.6651458515003323,
1144
+ "brier": 0.45323706188140617,
1145
+ "heldout": false,
1146
+ "domain": "reasoning",
1147
+ "ece": 0.1159557296037674,
1148
+ "proof_v2 (card)": NaN,
1149
+ "\u0394 vs proof_v2": NaN
1150
+ },
1151
+ {
1152
+ "task": "yelp/score",
1153
+ "n": 500,
1154
+ "acc": 0.668,
1155
+ "chance": 0.2,
1156
+ "nll": 0.7834262755662202,
1157
+ "brier": 0.4423740240858843,
1158
+ "heldout": false,
1159
+ "domain": "classification",
1160
+ "ece": 0.09068749487400055,
1161
+ "proof_v2 (card)": NaN,
1162
+ "\u0394 vs proof_v2": NaN
1163
+ }
1164
+ ],
1165
+ "head_to_head": [
1166
+ {
1167
+ "task": "ag_news/topic",
1168
+ "n": 120,
1169
+ "LightDec": 0.85,
1170
+ "proof_v2": 0.45,
1171
+ "\u0394": 0.39999999999999997
1172
+ },
1173
+ {
1174
+ "task": "agentharm/refuse",
1175
+ "n": 120,
1176
+ "LightDec": 0.6666666666666666,
1177
+ "proof_v2": 0.45,
1178
+ "\u0394": 0.21666666666666662
1179
+ },
1180
+ {
1181
+ "task": "agenttrek/finish_now",
1182
+ "n": 120,
1183
+ "LightDec": 0.7833333333333333,
1184
+ "proof_v2": 0.25833333333333336,
1185
+ "\u0394": 0.5249999999999999
1186
+ },
1187
+ {
1188
+ "task": "agenttrek/next_action_type",
1189
+ "n": 120,
1190
+ "LightDec": 0.875,
1191
+ "proof_v2": 0.075,
1192
+ "\u0394": 0.8
1193
+ },
1194
+ {
1195
+ "task": "anli/nli",
1196
+ "n": 120,
1197
+ "LightDec": 0.5916666666666667,
1198
+ "proof_v2": 0.325,
1199
+ "\u0394": 0.26666666666666666
1200
+ },
1201
+ {
1202
+ "task": "aqua_rat/math",
1203
+ "n": 120,
1204
+ "LightDec": 0.35,
1205
+ "proof_v2": 0.24166666666666667,
1206
+ "\u0394": 0.10833333333333331
1207
+ },
1208
+ {
1209
+ "task": "arc_challenge/mcq",
1210
+ "n": 120,
1211
+ "LightDec": 0.5333333333333333,
1212
+ "proof_v2": 0.30833333333333335,
1213
+ "\u0394": 0.22499999999999998
1214
+ },
1215
+ {
1216
+ "task": "arc_easy/mcq",
1217
+ "n": 120,
1218
+ "LightDec": 0.65,
1219
+ "proof_v2": 0.4583333333333333,
1220
+ "\u0394": 0.1916666666666667
1221
+ },
1222
+ {
1223
+ "task": "banking77/intent",
1224
+ "n": 120,
1225
+ "LightDec": 0.925,
1226
+ "proof_v2": 0.8666666666666667,
1227
+ "\u0394": 0.05833333333333335
1228
+ },
1229
+ {
1230
+ "task": "bigclonebench/clone",
1231
+ "n": 120,
1232
+ "LightDec": 0.9666666666666667,
1233
+ "proof_v2": 0.24166666666666667,
1234
+ "\u0394": 0.725
1235
+ },
1236
+ {
1237
+ "task": "bitext/category",
1238
+ "n": 120,
1239
+ "LightDec": 1.0,
1240
+ "proof_v2": 0.8583333333333333,
1241
+ "\u0394": 0.14166666666666672
1242
+ },
1243
+ {
1244
+ "task": "bitext/route",
1245
+ "n": 120,
1246
+ "LightDec": 1.0,
1247
+ "proof_v2": 0.925,
1248
+ "\u0394": 0.07499999999999996
1249
+ },
1250
+ {
1251
+ "task": "boolq/yes_no",
1252
+ "n": 120,
1253
+ "LightDec": 0.825,
1254
+ "proof_v2": 0.6666666666666666,
1255
+ "\u0394": 0.15833333333333333
1256
+ },
1257
+ {
1258
+ "task": "civil_comments/toxic",
1259
+ "n": 120,
1260
+ "LightDec": 0.9333333333333333,
1261
+ "proof_v2": 0.35833333333333334,
1262
+ "\u0394": 0.575
1263
+ },
1264
+ {
1265
+ "task": "clinc150/intent",
1266
+ "n": 120,
1267
+ "LightDec": 0.975,
1268
+ "proof_v2": 0.6416666666666667,
1269
+ "\u0394": 0.33333333333333326
1270
+ },
1271
+ {
1272
+ "task": "codexglue/code_to_doc",
1273
+ "n": 120,
1274
+ "LightDec": 1.0,
1275
+ "proof_v2": 0.9333333333333333,
1276
+ "\u0394": 0.06666666666666665
1277
+ },
1278
+ {
1279
+ "task": "codexglue/doc_to_code",
1280
+ "n": 120,
1281
+ "LightDec": 0.9833333333333333,
1282
+ "proof_v2": 0.975,
1283
+ "\u0394": 0.008333333333333304
1284
+ },
1285
+ {
1286
+ "task": "codexglue/func_name",
1287
+ "n": 120,
1288
+ "LightDec": 0.9916666666666667,
1289
+ "proof_v2": 0.9166666666666666,
1290
+ "\u0394": 0.07500000000000007
1291
+ },
1292
+ {
1293
+ "task": "codexglue/lang_id",
1294
+ "n": 120,
1295
+ "LightDec": 1.0,
1296
+ "proof_v2": 1.0,
1297
+ "\u0394": 0.0
1298
+ },
1299
+ {
1300
+ "task": "commonsense_qa/mcq",
1301
+ "n": 120,
1302
+ "LightDec": 0.6833333333333333,
1303
+ "proof_v2": 0.4083333333333333,
1304
+ "\u0394": 0.275
1305
+ },
1306
+ {
1307
+ "task": "counsel/critique_quality",
1308
+ "n": 120,
1309
+ "LightDec": 0.6083333333333333,
1310
+ "proof_v2": 0.25833333333333336,
1311
+ "\u0394": 0.3499999999999999
1312
+ },
1313
+ {
1314
+ "task": "counsel/step_has_error",
1315
+ "n": 120,
1316
+ "LightDec": 0.825,
1317
+ "proof_v2": 0.75,
1318
+ "\u0394": 0.07499999999999996
1319
+ },
1320
+ {
1321
+ "task": "devign/vulnerability",
1322
+ "n": 120,
1323
+ "LightDec": 0.6666666666666666,
1324
+ "proof_v2": 0.5666666666666667,
1325
+ "\u0394": 0.09999999999999998
1326
+ },
1327
+ {
1328
+ "task": "emotion/6way",
1329
+ "n": 120,
1330
+ "LightDec": 0.49166666666666664,
1331
+ "proof_v2": 0.36666666666666664,
1332
+ "\u0394": 0.125
1333
+ },
1334
+ {
1335
+ "task": "gsm8k/math",
1336
+ "n": 120,
1337
+ "LightDec": 0.7333333333333333,
1338
+ "proof_v2": 0.125,
1339
+ "\u0394": 0.6083333333333333
1340
+ },
1341
+ {
1342
+ "task": "hellaswag/continuation",
1343
+ "n": 120,
1344
+ "LightDec": 0.6083333333333333,
1345
+ "proof_v2": 0.2916666666666667,
1346
+ "\u0394": 0.3166666666666666
1347
+ },
1348
+ {
1349
+ "task": "hotpotqa/comparison_yes_no",
1350
+ "n": 26,
1351
+ "LightDec": 0.9230769230769231,
1352
+ "proof_v2": 0.46153846153846156,
1353
+ "\u0394": 0.46153846153846156
1354
+ },
1355
+ {
1356
+ "task": "hotpotqa/retrieve",
1357
+ "n": 120,
1358
+ "LightDec": 0.875,
1359
+ "proof_v2": 0.2916666666666667,
1360
+ "\u0394": 0.5833333333333333
1361
+ },
1362
+ {
1363
+ "task": "humaneval/completion",
1364
+ "n": 119,
1365
+ "LightDec": 0.8487394957983193,
1366
+ "proof_v2": 0.5126050420168067,
1367
+ "\u0394": 0.33613445378151263
1368
+ },
1369
+ {
1370
+ "task": "jailbreak/detect",
1371
+ "n": 120,
1372
+ "LightDec": 0.9833333333333333,
1373
+ "proof_v2": 0.5666666666666667,
1374
+ "\u0394": 0.41666666666666663
1375
+ },
1376
+ {
1377
+ "task": "long/contract_clause",
1378
+ "n": 120,
1379
+ "LightDec": 1.0,
1380
+ "proof_v2": 0.23333333333333334,
1381
+ "\u0394": 0.7666666666666666
1382
+ },
1383
+ {
1384
+ "task": "long/email_thread",
1385
+ "n": 120,
1386
+ "LightDec": 1.0,
1387
+ "proof_v2": 0.05,
1388
+ "\u0394": 0.95
1389
+ },
1390
+ {
1391
+ "task": "long/service_log",
1392
+ "n": 120,
1393
+ "LightDec": 1.0,
1394
+ "proof_v2": 0.24166666666666667,
1395
+ "\u0394": 0.7583333333333333
1396
+ },
1397
+ {
1398
+ "task": "massive_en/intent",
1399
+ "n": 120,
1400
+ "LightDec": 0.95,
1401
+ "proof_v2": 0.7416666666666667,
1402
+ "\u0394": 0.20833333333333326
1403
+ },
1404
+ {
1405
+ "task": "mbpp/bugspot",
1406
+ "n": 120,
1407
+ "LightDec": 0.8583333333333333,
1408
+ "proof_v2": 0.575,
1409
+ "\u0394": 0.2833333333333333
1410
+ },
1411
+ {
1412
+ "task": "mbpp/solution",
1413
+ "n": 120,
1414
+ "LightDec": 1.0,
1415
+ "proof_v2": 0.85,
1416
+ "\u0394": 0.15000000000000002
1417
+ },
1418
+ {
1419
+ "task": "mmlu/mcq",
1420
+ "n": 120,
1421
+ "LightDec": 0.4083333333333333,
1422
+ "proof_v2": 0.30833333333333335,
1423
+ "\u0394": 0.09999999999999998
1424
+ },
1425
+ {
1426
+ "task": "mnli/claim",
1427
+ "n": 120,
1428
+ "LightDec": 0.875,
1429
+ "proof_v2": 0.31666666666666665,
1430
+ "\u0394": 0.5583333333333333
1431
+ },
1432
+ {
1433
+ "task": "openbookqa/mcq",
1434
+ "n": 120,
1435
+ "LightDec": 0.65,
1436
+ "proof_v2": 0.2833333333333333,
1437
+ "\u0394": 0.3666666666666667
1438
+ },
1439
+ {
1440
+ "task": "policy/access_control_transfer",
1441
+ "n": 120,
1442
+ "LightDec": 1.0,
1443
+ "proof_v2": 0.23333333333333334,
1444
+ "\u0394": 0.7666666666666666
1445
+ },
1446
+ {
1447
+ "task": "policy/count_threshold_transfer",
1448
+ "n": 120,
1449
+ "LightDec": 0.8666666666666667,
1450
+ "proof_v2": 0.10833333333333334,
1451
+ "\u0394": 0.7583333333333333
1452
+ },
1453
+ {
1454
+ "task": "policy/free_shipping_transfer",
1455
+ "n": 120,
1456
+ "LightDec": 0.925,
1457
+ "proof_v2": 0.7083333333333334,
1458
+ "\u0394": 0.21666666666666667
1459
+ },
1460
+ {
1461
+ "task": "policy/invoice_overdue_transfer",
1462
+ "n": 120,
1463
+ "LightDec": 0.8083333333333333,
1464
+ "proof_v2": 0.6583333333333333,
1465
+ "\u0394": 0.15000000000000002
1466
+ },
1467
+ {
1468
+ "task": "policy/invoice_total_transfer",
1469
+ "n": 120,
1470
+ "LightDec": 0.45,
1471
+ "proof_v2": 0.525,
1472
+ "\u0394": -0.07500000000000001
1473
+ },
1474
+ {
1475
+ "task": "policy/refund_approval_transfer",
1476
+ "n": 120,
1477
+ "LightDec": 0.9583333333333334,
1478
+ "proof_v2": 0.325,
1479
+ "\u0394": 0.6333333333333333
1480
+ },
1481
+ {
1482
+ "task": "policy/return_window_transfer",
1483
+ "n": 120,
1484
+ "LightDec": 1.0,
1485
+ "proof_v2": 0.3416666666666667,
1486
+ "\u0394": 0.6583333333333333
1487
+ },
1488
+ {
1489
+ "task": "policy/sla_urgency_transfer",
1490
+ "n": 120,
1491
+ "LightDec": 0.5166666666666667,
1492
+ "proof_v2": 0.44166666666666665,
1493
+ "\u0394": 0.07500000000000007
1494
+ },
1495
+ {
1496
+ "task": "policy/table_compare_transfer",
1497
+ "n": 120,
1498
+ "LightDec": 0.9166666666666666,
1499
+ "proof_v2": 0.49166666666666664,
1500
+ "\u0394": 0.425
1501
+ },
1502
+ {
1503
+ "task": "policy/table_count_transfer",
1504
+ "n": 120,
1505
+ "LightDec": 0.31666666666666665,
1506
+ "proof_v2": 0.15833333333333333,
1507
+ "\u0394": 0.15833333333333333
1508
+ },
1509
+ {
1510
+ "task": "policy/table_extreme_transfer",
1511
+ "n": 120,
1512
+ "LightDec": 0.925,
1513
+ "proof_v2": 0.14166666666666666,
1514
+ "\u0394": 0.7833333333333334
1515
+ },
1516
+ {
1517
+ "task": "prompt_injections/detect",
1518
+ "n": 116,
1519
+ "LightDec": 0.5948275862068966,
1520
+ "proof_v2": 0.5431034482758621,
1521
+ "\u0394": 0.051724137931034475
1522
+ },
1523
+ {
1524
+ "task": "qasc/mcq",
1525
+ "n": 120,
1526
+ "LightDec": 0.9916666666666667,
1527
+ "proof_v2": 0.8583333333333333,
1528
+ "\u0394": 0.13333333333333341
1529
+ },
1530
+ {
1531
+ "task": "sciq/mcq",
1532
+ "n": 120,
1533
+ "LightDec": 0.9666666666666667,
1534
+ "proof_v2": 0.8666666666666667,
1535
+ "\u0394": 0.09999999999999998
1536
+ },
1537
+ {
1538
+ "task": "scitail/support",
1539
+ "n": 120,
1540
+ "LightDec": 0.9583333333333334,
1541
+ "proof_v2": 0.425,
1542
+ "\u0394": 0.5333333333333334
1543
+ },
1544
+ {
1545
+ "task": "snli/contradicts",
1546
+ "n": 120,
1547
+ "LightDec": 1.0,
1548
+ "proof_v2": 0.35833333333333334,
1549
+ "\u0394": 0.6416666666666666
1550
+ },
1551
+ {
1552
+ "task": "snli/must_be_true",
1553
+ "n": 120,
1554
+ "LightDec": 0.9916666666666667,
1555
+ "proof_v2": 0.925,
1556
+ "\u0394": 0.06666666666666665
1557
+ },
1558
+ {
1559
+ "task": "snli/nli",
1560
+ "n": 120,
1561
+ "LightDec": 0.9083333333333333,
1562
+ "proof_v2": 0.3416666666666667,
1563
+ "\u0394": 0.5666666666666667
1564
+ },
1565
+ {
1566
+ "task": "sst5/score",
1567
+ "n": 120,
1568
+ "LightDec": 0.38333333333333336,
1569
+ "proof_v2": 0.2833333333333333,
1570
+ "\u0394": 0.10000000000000003
1571
+ },
1572
+ {
1573
+ "task": "tev1_test/ag_news",
1574
+ "n": 120,
1575
+ "LightDec": 0.925,
1576
+ "proof_v2": 0.425,
1577
+ "\u0394": 0.5
1578
+ },
1579
+ {
1580
+ "task": "tev1_test/banking77",
1581
+ "n": 120,
1582
+ "LightDec": 0.7166666666666667,
1583
+ "proof_v2": 0.55,
1584
+ "\u0394": 0.16666666666666663
1585
+ },
1586
+ {
1587
+ "task": "tev1_test/boolq",
1588
+ "n": 120,
1589
+ "LightDec": 0.8416666666666667,
1590
+ "proof_v2": 0.5583333333333333,
1591
+ "\u0394": 0.2833333333333333
1592
+ },
1593
+ {
1594
+ "task": "tev1_test/mnli",
1595
+ "n": 120,
1596
+ "LightDec": 0.75,
1597
+ "proof_v2": 0.39166666666666666,
1598
+ "\u0394": 0.35833333333333334
1599
+ },
1600
+ {
1601
+ "task": "tev1_test/policy",
1602
+ "n": 120,
1603
+ "LightDec": 0.5,
1604
+ "proof_v2": 0.4083333333333333,
1605
+ "\u0394": 0.09166666666666667
1606
+ },
1607
+ {
1608
+ "task": "tev1_test/routing",
1609
+ "n": 120,
1610
+ "LightDec": 0.44166666666666665,
1611
+ "proof_v2": 0.2916666666666667,
1612
+ "\u0394": 0.14999999999999997
1613
+ },
1614
+ {
1615
+ "task": "tev1_test/sst5",
1616
+ "n": 120,
1617
+ "LightDec": 0.39166666666666666,
1618
+ "proof_v2": 0.275,
1619
+ "\u0394": 0.11666666666666664
1620
+ },
1621
+ {
1622
+ "task": "triage/support_email",
1623
+ "n": 120,
1624
+ "LightDec": 0.9583333333333334,
1625
+ "proof_v2": 0.39166666666666666,
1626
+ "\u0394": 0.5666666666666667
1627
+ },
1628
+ {
1629
+ "task": "typed_decisions/agent_trace_observability",
1630
+ "n": 120,
1631
+ "LightDec": 0.8083333333333333,
1632
+ "proof_v2": 0.3333333333333333,
1633
+ "\u0394": 0.47500000000000003
1634
+ },
1635
+ {
1636
+ "task": "typed_decisions/customer_service",
1637
+ "n": 120,
1638
+ "LightDec": 0.7416666666666667,
1639
+ "proof_v2": 0.2833333333333333,
1640
+ "\u0394": 0.45833333333333337
1641
+ },
1642
+ {
1643
+ "task": "typed_decisions/invoice_processing",
1644
+ "n": 120,
1645
+ "LightDec": 0.8166666666666667,
1646
+ "proof_v2": 0.475,
1647
+ "\u0394": 0.3416666666666667
1648
+ },
1649
+ {
1650
+ "task": "typed_decisions/security_incidents",
1651
+ "n": 120,
1652
+ "LightDec": 0.7583333333333333,
1653
+ "proof_v2": 0.5,
1654
+ "\u0394": 0.2583333333333333
1655
+ },
1656
+ {
1657
+ "task": "winogrande/blank",
1658
+ "n": 120,
1659
+ "LightDec": 0.6916666666666667,
1660
+ "proof_v2": 0.5583333333333333,
1661
+ "\u0394": 0.1333333333333333
1662
+ },
1663
+ {
1664
+ "task": "yelp/score",
1665
+ "n": 120,
1666
+ "LightDec": 0.625,
1667
+ "proof_v2": 0.325,
1668
+ "\u0394": 0.3
1669
+ }
1670
+ ],
1671
+ "head_to_head_error": null,
1672
+ "v3_benchmarks": {
1673
+ "tev1": {
1674
+ "per_task": [
1675
+ {
1676
+ "task": "tev1_test/ag_news",
1677
+ "n": 150,
1678
+ "acc": 0.92,
1679
+ "overlaps_LightDec_training_source": true
1680
+ },
1681
+ {
1682
+ "task": "tev1_test/banking77",
1683
+ "n": 200,
1684
+ "acc": 0.72,
1685
+ "overlaps_LightDec_training_source": false
1686
+ },
1687
+ {
1688
+ "task": "tev1_test/boolq",
1689
+ "n": 200,
1690
+ "acc": 0.835,
1691
+ "overlaps_LightDec_training_source": true
1692
+ },
1693
+ {
1694
+ "task": "tev1_test/mnli",
1695
+ "n": 300,
1696
+ "acc": 0.7233333333333334,
1697
+ "overlaps_LightDec_training_source": true
1698
+ },
1699
+ {
1700
+ "task": "tev1_test/policy",
1701
+ "n": 1200,
1702
+ "acc": 0.5291666666666667,
1703
+ "overlaps_LightDec_training_source": false
1704
+ },
1705
+ {
1706
+ "task": "tev1_test/routing",
1707
+ "n": 600,
1708
+ "acc": 0.4583333333333333,
1709
+ "overlaps_LightDec_training_source": false
1710
+ },
1711
+ {
1712
+ "task": "tev1_test/sst5",
1713
+ "n": 150,
1714
+ "acc": 0.3933333333333333,
1715
+ "overlaps_LightDec_training_source": false
1716
+ }
1717
+ ],
1718
+ "all": 0.5839285714285715,
1719
+ "no_overlap": 0.5176744186046511,
1720
+ "n": 2800,
1721
+ "published_tev1_4b": {
1722
+ "main decisions (1,000)": 0.88,
1723
+ "policy transfer (300)": 1.0
1724
+ }
1725
+ },
1726
+ "triage": {
1727
+ "emails": 101,
1728
+ "per_question": {
1729
+ "topic": 0.9900990099009901,
1730
+ "refund": 1.0,
1731
+ "breakage": 0.9900990099009901,
1732
+ "anger": 0.9207920792079208,
1733
+ "judgment": 0.8217821782178217
1734
+ },
1735
+ "pile_accuracy": 0.9405940594059405
1736
+ },
1737
+ "long_context": {
1738
+ "max_len": 2048,
1739
+ "per_task": [
1740
+ {
1741
+ "task": "long/contract_clause",
1742
+ "n": 150,
1743
+ "acc": 1.0,
1744
+ "chance": 0.19999999999999996
1745
+ },
1746
+ {
1747
+ "task": "long/email_thread",
1748
+ "n": 150,
1749
+ "acc": 1.0,
1750
+ "chance": 0.25
1751
+ },
1752
+ {
1753
+ "task": "long/service_log",
1754
+ "n": 150,
1755
+ "acc": 1.0,
1756
+ "chance": 0.19999999999999996
1757
+ }
1758
+ ]
1759
+ }
1760
+ },
1761
+ "architecture": "FalconDec",
1762
+ "latency": {
1763
+ "gpu": "NVIDIA RTX PRO 6000 Blackwell Server Edition",
1764
+ "calls": [
1765
+ {
1766
+ "device": "cuda",
1767
+ "weights": "fp16",
1768
+ "questions": 1,
1769
+ "p50_ms": 10.05,
1770
+ "p95_ms": 10.43
1771
+ },
1772
+ {
1773
+ "device": "cuda",
1774
+ "weights": "fp16",
1775
+ "questions": 5,
1776
+ "p50_ms": 11.32,
1777
+ "p95_ms": 11.41
1778
+ },
1779
+ {
1780
+ "device": "cuda",
1781
+ "weights": "fp16",
1782
+ "questions": 10,
1783
+ "p50_ms": 12.57,
1784
+ "p95_ms": 12.69
1785
+ },
1786
+ {
1787
+ "device": "cpu fp32",
1788
+ "weights": "int8 file",
1789
+ "questions": 1,
1790
+ "p50_ms": 48.4,
1791
+ "p95_ms": 48.8
1792
+ }
1793
+ ],
1794
+ "throughput_per_s": 2588.8502560238694
1795
+ },
1796
+ "env": {
1797
+ "torch": "2.9.0+cu130",
1798
+ "transformers": "5.17.0",
1799
+ "python": "3.12.12",
1800
+ "gpu": "NVIDIA RTX PRO 6000 Blackwell Server Edition"
1801
+ }
1802
+ }
compact-int8/model_int8.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a93d550e1c95773af46f11eefcde831edd723e694b20d52e127c1b790ff38959
3
+ size 160530522
compact-int8/tokenizer/tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
compact-int8/tokenizer/tokenizer_config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": true,
4
+ "cls_token": "[CLS]",
5
+ "is_local": false,
6
+ "local_files_only": false,
7
+ "mask_token": "[MASK]",
8
+ "model_input_names": [
9
+ "input_ids",
10
+ "attention_mask"
11
+ ],
12
+ "model_max_length": 8192,
13
+ "pad_token": "[PAD]",
14
+ "sep_token": "[SEP]",
15
+ "tokenizer_class": "TokenizersBackend",
16
+ "unk_token": "[UNK]"
17
+ }
encoder/config.json ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "ModernBertForMaskedLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 50281,
8
+ "causal_mask": false,
9
+ "classifier_activation": "gelu",
10
+ "classifier_bias": false,
11
+ "classifier_dropout": 0.0,
12
+ "classifier_pooling": "mean",
13
+ "cls_token_id": 50281,
14
+ "decoder_bias": true,
15
+ "deterministic_flash_attn": false,
16
+ "dtype": "float32",
17
+ "embedding_dropout": 0.0,
18
+ "eos_token_id": 50282,
19
+ "global_attn_every_n_layers": 3,
20
+ "gradient_checkpointing": false,
21
+ "hidden_activation": "gelu",
22
+ "hidden_size": 768,
23
+ "initializer_cutoff_factor": 2.0,
24
+ "initializer_range": 0.02,
25
+ "intermediate_size": 1152,
26
+ "is_causal": false,
27
+ "layer_norm_eps": 1e-05,
28
+ "layer_types": [
29
+ "full_attention",
30
+ "sliding_attention",
31
+ "sliding_attention",
32
+ "full_attention",
33
+ "sliding_attention",
34
+ "sliding_attention",
35
+ "full_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "full_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "full_attention",
42
+ "sliding_attention",
43
+ "sliding_attention",
44
+ "full_attention",
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "full_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "full_attention"
51
+ ],
52
+ "local_attention": 128,
53
+ "max_position_embeddings": 7999,
54
+ "mlp_bias": false,
55
+ "mlp_dropout": 0.0,
56
+ "model_type": "modernbert",
57
+ "norm_bias": false,
58
+ "norm_eps": 1e-05,
59
+ "num_attention_heads": 12,
60
+ "num_hidden_layers": 22,
61
+ "pad_token_id": 50283,
62
+ "position_embedding_type": "sans_pos",
63
+ "rope_parameters": {
64
+ "full_attention": {
65
+ "rope_theta": 160000.0,
66
+ "rope_type": "default"
67
+ },
68
+ "sliding_attention": {
69
+ "rope_theta": 160000.0,
70
+ "rope_type": "default"
71
+ }
72
+ },
73
+ "sep_token_id": 50282,
74
+ "sparse_pred_ignore_index": -100,
75
+ "sparse_prediction": false,
76
+ "tie_word_embeddings": true,
77
+ "transformers_version": "5.17.0",
78
+ "vocab_size": 50368
79
+ }
falcondec_config.json ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "max_len": 2048,
3
+ "long_max_len": 2048,
4
+ "long_opts_threshold": 24,
5
+ "head_max_len": 192,
6
+ "max_tok_per_opt": 24,
7
+ "max_opts_single_pass": 96,
8
+ "interact_layers": 2,
9
+ "interact_heads": 8,
10
+ "name": "LightDec_V2_Long",
11
+ "backbone": "jhu-clsp/ettin-encoder-150m",
12
+ "backbone_init": "pretrained",
13
+ "lineage": [
14
+ "jhu-clsp/ettin-encoder-150m"
15
+ ],
16
+ "parent_version": null,
17
+ "special": {
18
+ "cls": 50281,
19
+ "sep": 50282,
20
+ "mask": 50284,
21
+ "pad": 50283
22
+ },
23
+ "defer_threshold": 0.7,
24
+ "qtypes": {
25
+ "choice": 0,
26
+ "noul": 1,
27
+ "score": 2
28
+ },
29
+ "n_buckets": 4,
30
+ "version": "1.0.0",
31
+ "notebook_version": "V3.7",
32
+ "created": "2026-09-28T10:52:36.426087Z",
33
+ "temperature": [
34
+ [
35
+ 2.37591814994812,
36
+ 2.0426506996154785,
37
+ 1.6455482244491577,
38
+ 1.8417285680770874
39
+ ],
40
+ [
41
+ 2.4885058403015137,
42
+ 2.4885058403015137,
43
+ 2.4885058403015137,
44
+ 2.4885058403015137
45
+ ],
46
+ [
47
+ 2.2925639152526855,
48
+ 2.2925639152526855,
49
+ 2.2925639152526855,
50
+ 2.2925639152526855
51
+ ]
52
+ ],
53
+ "weights": "model.safetensors"
54
+ }
falcondec_modeling.py ADDED
@@ -0,0 +1,358 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """FalconDec — Falcon Decision model: single-pass, typed, calibrated closed-set decisions.
3
+
4
+ Layout : [CLS] question [SEP] [MASK] opt_1 ... [MASK] opt_k [SEP] state [SEP]
5
+ Head : marker vectors + CLS context + question-type embedding
6
+ -> permutation-equivariant set transformer (options attend to each other; no positions)
7
+ -> MLP -> one logit per option -> softmax over this question's options
8
+ Calib. : temperature per (question type, option-count bucket), stored in the checkpoint
9
+ """
10
+ from __future__ import annotations
11
+
12
+ import contextlib
13
+ import json
14
+ import shutil
15
+ from pathlib import Path
16
+
17
+ import numpy as np
18
+ import torch
19
+ import torch.nn as nn
20
+
21
+ QTYPES = {"choice": 0, "noul": 1, "score": 2}
22
+ N_BUCKETS = 4
23
+ CONFIG_FILE = "falcondec_config.json"
24
+ FP16_FILE = "model.safetensors"
25
+ INT8_FILE = "model_int8.safetensors"
26
+ MODEL_KEYS = ("input_ids", "attention_mask", "marker_pos", "marker_mask", "qtype")
27
+
28
+
29
+ def n_bucket(n: int) -> int:
30
+ return 0 if n <= 2 else 1 if n <= 5 else 2 if n <= 12 else 3
31
+
32
+
33
+ def special_ids(tok) -> dict:
34
+ sp = {"cls": tok.cls_token_id, "sep": tok.sep_token_id, "mask": tok.mask_token_id, "pad": tok.pad_token_id}
35
+ if sp["cls"] is None:
36
+ sp["cls"] = tok.bos_token_id
37
+ if sp["sep"] is None:
38
+ sp["sep"] = tok.eos_token_id
39
+ if sp["pad"] is None:
40
+ sp["pad"] = sp["sep"]
41
+ if sp["mask"] is None:
42
+ raise ValueError("FalconDec needs a tokenizer with a mask token (used as the option marker).")
43
+ return {k: int(v) for k, v in sp.items()}
44
+
45
+
46
+ def _as_list(x):
47
+ return x.tolist() if hasattr(x, "tolist") else list(x)
48
+
49
+
50
+ def assemble(q_ids, opt_ids, s_ids, sp, max_len=512, head_max_len=192, max_tok_per_opt=24,
51
+ long_max_len=2048, long_opts_threshold=24):
52
+ """Build one input sequence. Returns (ids, marker_positions)."""
53
+ n = len(opt_ids)
54
+ eff = max_len if n <= long_opts_threshold else max(max_len, long_max_len)
55
+ q = _as_list(q_ids)[:96]
56
+ need = len(q) + 2 + n * (max_tok_per_opt + 1)
57
+ head = min(eff - 64, max(head_max_len, need))
58
+ per = max(2, min(max_tok_per_opt, (head - len(q) - 2) // max(n, 1) - 1))
59
+ ids = [sp["cls"]] + q + [sp["sep"]]
60
+ markers = []
61
+ for o in opt_ids:
62
+ markers.append(len(ids))
63
+ ids.append(sp["mask"])
64
+ ids.extend(_as_list(o[:per]))
65
+ ids.append(sp["sep"])
66
+ room = eff - len(ids) - 1
67
+ if room > 0 and s_ids is not None and len(s_ids) > 0:
68
+ ids.extend(_as_list(s_ids[:room]))
69
+ ids.append(sp["sep"])
70
+ if len(ids) > eff:
71
+ ids = ids[:eff]
72
+ if markers and markers[-1] >= len(ids):
73
+ raise ValueError("Too many options for one pass; reduce options or raise long_max_len.")
74
+ return ids, markers
75
+
76
+
77
+ def collate_features(feats, pad_id, device=None):
78
+ """feats: list of (ids, markers, qtype_index) -> dict of padded tensors matching FalconDec.forward."""
79
+ B = len(feats)
80
+ T = max(len(f[0]) for f in feats)
81
+ K = max(len(f[1]) for f in feats)
82
+ ids = torch.full((B, T), pad_id, dtype=torch.long)
83
+ att = torch.zeros((B, T), dtype=torch.long)
84
+ mpos = torch.zeros((B, K), dtype=torch.long)
85
+ mmask = torch.zeros((B, K), dtype=torch.bool)
86
+ qt = torch.zeros(B, dtype=torch.long)
87
+ for i, (x, mk, q) in enumerate(feats):
88
+ ids[i, : len(x)] = torch.as_tensor(x, dtype=torch.long)
89
+ att[i, : len(x)] = 1
90
+ mpos[i, : len(mk)] = torch.as_tensor(mk, dtype=torch.long)
91
+ mmask[i, : len(mk)] = True
92
+ qt[i] = int(q)
93
+ out = dict(input_ids=ids, attention_mask=att, marker_pos=mpos, marker_mask=mmask, qtype=qt)
94
+ if device is not None:
95
+ out = {k: v.to(device, non_blocking=True) for k, v in out.items()}
96
+ return out
97
+
98
+
99
+ class OptionInteraction(nn.Module):
100
+ """Set transformer over the options of one question (no positional encoding => order-equivariant)."""
101
+
102
+ def __init__(self, d, n_layers=2, n_heads=8, dropout=0.1):
103
+ super().__init__()
104
+ layer = nn.TransformerEncoderLayer(d, n_heads, dim_feedforward=2 * d, dropout=dropout,
105
+ activation="gelu", batch_first=True, norm_first=True)
106
+ self.enc = nn.TransformerEncoder(layer, n_layers, enable_nested_tensor=False)
107
+
108
+ def forward(self, x, mask):
109
+ return self.enc(x, src_key_padding_mask=~mask)
110
+
111
+
112
+ class FalconDec(nn.Module):
113
+ def __init__(self, encoder, fcfg: dict):
114
+ super().__init__()
115
+ self.encoder = encoder
116
+ self.fcfg = dict(fcfg)
117
+ d = encoder.config.hidden_size
118
+ self.qtype_emb = nn.Embedding(len(QTYPES), d)
119
+ self.ctx_proj = nn.Linear(d, d)
120
+ self.opt_norm = nn.LayerNorm(d)
121
+ self.interact = OptionInteraction(d, int(self.fcfg.get("interact_layers", 2)),
122
+ int(self.fcfg.get("interact_heads", 8)))
123
+ self.scorer = nn.Sequential(nn.Linear(d, d), nn.GELU(), nn.Dropout(0.1), nn.Linear(d, 1))
124
+ self.register_buffer("temperature", torch.ones(len(QTYPES), N_BUCKETS))
125
+
126
+ @property
127
+ def device(self):
128
+ return next(self.parameters()).device
129
+
130
+ def num_parameters(self):
131
+ return sum(p.numel() for p in self.parameters())
132
+
133
+ def forward(self, input_ids, attention_mask, marker_pos, marker_mask, qtype):
134
+ h = self.encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state
135
+ B, K = marker_pos.shape
136
+ idx = marker_pos.unsqueeze(-1).expand(B, K, h.size(-1))
137
+ opt = h.gather(1, idx)
138
+ ctx = self.ctx_proj(h[:, 0]).unsqueeze(1)
139
+ x = self.opt_norm(opt + ctx + self.qtype_emb(qtype).unsqueeze(1))
140
+ x = self.interact(x, marker_mask)
141
+ logits = self.scorer(x).squeeze(-1).float()
142
+ return logits.masked_fill(~marker_mask, -1e4)
143
+
144
+
145
+ # ------------------------------------------------------------------ storage
146
+ def clean_state_dict(sd):
147
+ return {k.replace("_orig_mod.", ""): v.detach().cpu().contiguous() for k, v in sd.items()}
148
+
149
+
150
+ def quantize_int8(sd, min_numel=4096):
151
+ """Per-output-channel symmetric int8 for every matrix; fp16 for everything else."""
152
+ out = {}
153
+ for k, v in sd.items():
154
+ if v.is_floating_point() and v.ndim == 2 and v.numel() >= min_numel:
155
+ w = v.float()
156
+ s = (w.abs().amax(dim=1, keepdim=True) / 127.0).clamp_min(1e-12)
157
+ out[k] = torch.round(w / s).clamp_(-127, 127).to(torch.int8).contiguous()
158
+ out[k + "::scale"] = s.contiguous()
159
+ elif v.is_floating_point():
160
+ out[k] = v.half().contiguous()
161
+ else:
162
+ out[k] = v.contiguous()
163
+ return out
164
+
165
+
166
+ def dequantize_int8(sd):
167
+ out = {}
168
+ for k, v in sd.items():
169
+ if k.endswith("::scale"):
170
+ continue
171
+ s = sd.get(k + "::scale")
172
+ out[k] = (v.float() * s).half() if s is not None else v
173
+ return out
174
+
175
+
176
+ def save_falcondec(model, tok, out_dir, int8=False, extra_files=None):
177
+ from safetensors.torch import save_file
178
+ out = Path(out_dir)
179
+ out.mkdir(parents=True, exist_ok=True)
180
+ enc_cfg = model.encoder.config
181
+ if hasattr(enc_cfg, "reference_compile"):
182
+ enc_cfg.reference_compile = False
183
+ enc_cfg.save_pretrained(str(out / "encoder"))
184
+ tok.save_pretrained(str(out / "tokenizer"))
185
+ sd = clean_state_dict(model.state_dict())
186
+ fc = dict(model.fcfg)
187
+ fc["temperature"] = model.temperature.detach().float().cpu().tolist()
188
+ fc["weights"] = INT8_FILE if int8 else FP16_FILE
189
+ if int8:
190
+ save_file(quantize_int8(sd), str(out / INT8_FILE), metadata={"format": "falcondec-int8"})
191
+ else:
192
+ save_file({k: (v.half() if v.is_floating_point() else v) for k, v in sd.items()},
193
+ str(out / FP16_FILE), metadata={"format": "falcondec-fp16"})
194
+ (out / CONFIG_FILE).write_text(json.dumps(fc, indent=2, default=str), encoding="utf-8")
195
+ try:
196
+ shutil.copy(__file__, out / "falcondec_modeling.py")
197
+ except Exception:
198
+ pass
199
+ for name, content in (extra_files or {}).items():
200
+ (out / name).write_text(content, encoding="utf-8")
201
+ return out
202
+
203
+
204
+ def load_falcondec(path, device=None, dtype=None, attn_implementation="sdpa"):
205
+ """Load a FalconDec directory (fp16 or int8) or Hub repo. Returns (model, tokenizer)."""
206
+ from safetensors.torch import load_file
207
+ from transformers import AutoConfig, AutoModel, AutoTokenizer
208
+ p = Path(path)
209
+ if not p.exists():
210
+ from huggingface_hub import snapshot_download
211
+ p = Path(snapshot_download(str(path)))
212
+ fc = json.loads((p / CONFIG_FILE).read_text(encoding="utf-8"))
213
+ ecfg = AutoConfig.from_pretrained(str(p / "encoder"))
214
+ if hasattr(ecfg, "reference_compile"):
215
+ ecfg.reference_compile = False
216
+ try:
217
+ enc = AutoModel.from_config(ecfg, attn_implementation=attn_implementation)
218
+ except Exception:
219
+ enc = AutoModel.from_config(ecfg)
220
+ model = FalconDec(enc, fc)
221
+ wf = p / fc.get("weights", FP16_FILE)
222
+ if not wf.exists():
223
+ wf = p / (INT8_FILE if (p / INT8_FILE).exists() else FP16_FILE)
224
+ sd = load_file(str(wf))
225
+ if wf.name == INT8_FILE:
226
+ sd = dequantize_int8(sd)
227
+ missing, unexpected = model.load_state_dict(sd, strict=False)
228
+ if missing or unexpected:
229
+ print(f"[FalconDec] load warning: missing={list(missing)[:5]} unexpected={list(unexpected)[:5]}")
230
+ tok = AutoTokenizer.from_pretrained(str(p / "tokenizer"))
231
+ dev = torch.device(device) if device is not None else torch.device("cuda" if torch.cuda.is_available() else "cpu")
232
+ model.to(dev)
233
+ if dtype is not None:
234
+ model.to(dtype)
235
+ model.eval()
236
+ return model, tok
237
+
238
+
239
+ # ------------------------------------------------------------------ inference
240
+ def _amp(model):
241
+ if model.device.type == "cuda" and next(model.parameters()).dtype == torch.float32:
242
+ dt = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
243
+ return torch.autocast("cuda", dtype=dt)
244
+ return contextlib.nullcontext()
245
+
246
+
247
+ def _state_text(state):
248
+ if state is None:
249
+ return ""
250
+ return state if isinstance(state, str) else json.dumps(state, ensure_ascii=False)
251
+
252
+
253
+ @torch.no_grad()
254
+ def score_items(model, tok, items, batch_size=32):
255
+ """items: [{"state", "question", "options", "type"?, "option_tokens"?, "seq_len"?}] -> list of prob arrays."""
256
+ fc = model.fcfg
257
+ M = int(fc.get("max_opts_single_pass", 96))
258
+ out = [None] * len(items)
259
+ small = [i for i, it in enumerate(items) if len(it["options"]) <= M]
260
+ for i, it in enumerate(items):
261
+ if len(it["options"]) > M:
262
+ out[i] = _score_large(model, tok, it, batch_size)
263
+ if not small:
264
+ return out
265
+ sp = fc["special"]
266
+ prepared = {}
267
+ for i in small:
268
+ it = items[i]
269
+ if len(it["options"]) < 2:
270
+ raise ValueError("Every question needs at least two options.")
271
+ budget = int(it.get("option_tokens", fc["max_tok_per_opt"]))
272
+ q_ids = tok(it.get("question", ""), add_special_tokens=False, truncation=True, max_length=96)["input_ids"]
273
+ o_ids = tok([str(o) for o in it["options"]], add_special_tokens=False, truncation=True,
274
+ max_length=budget)["input_ids"]
275
+ s = _state_text(it.get("state"))
276
+ s_ids = tok(s, add_special_tokens=False, truncation=True,
277
+ max_length=int(fc["long_max_len"]))["input_ids"] if s else []
278
+ ids, mk = assemble(q_ids, o_ids, s_ids, sp, int(it.get("seq_len", fc["max_len"])), fc["head_max_len"],
279
+ budget, fc["long_max_len"], fc["long_opts_threshold"])
280
+ prepared[i] = (ids, mk, QTYPES[it.get("type", "choice")])
281
+ order = sorted(small, key=lambda i: len(prepared[i][0]))
282
+ model.eval()
283
+ for b0 in range(0, len(order), batch_size):
284
+ idxs = order[b0: b0 + batch_size]
285
+ batch = collate_features([prepared[i] for i in idxs], sp["pad"], model.device)
286
+ with _amp(model):
287
+ logits = model(**batch)
288
+ for j, i in enumerate(idxs):
289
+ n = len(prepared[i][1])
290
+ T = model.temperature[prepared[i][2], n_bucket(n)].float().clamp_min(1e-3)
291
+ out[i] = torch.softmax(logits[j, :n].float() / T, -1).cpu().numpy()
292
+ return out
293
+
294
+
295
+ def _score_large(model, tok, it, batch_size):
296
+ """Tournament for > max_opts_single_pass options: keep the best of each chunk, then one final pass."""
297
+ M = int(model.fcfg.get("max_opts_single_pass", 96))
298
+ opts = list(it["options"])
299
+ cand = list(range(len(opts)))
300
+ while len(cand) > M:
301
+ groups = [cand[i: i + M] for i in range(0, len(cand), M)]
302
+ keep = max(1, M // len(groups))
303
+ ps = score_items(model, tok, [dict(it, options=[opts[c] for c in g]) for g in groups], batch_size)
304
+ cand = [g[t] for g, p in zip(groups, ps) for t in np.argsort(-p)[:keep]]
305
+ p = score_items(model, tok, [dict(it, options=[opts[c] for c in cand])], batch_size)[0]
306
+ full = np.zeros(len(opts), dtype=np.float32)
307
+ full[cand] = p
308
+ return full
309
+
310
+
311
+ def _normalize_question(q):
312
+ qtype = q.get("type", "choice")
313
+ text = q.get("question") or q.get("instructions") or ""
314
+ if qtype == "noul":
315
+ lab = q.get("labels") or {}
316
+ return qtype, text, [True, False], [str(lab.get("true", "Yes")), str(lab.get("false", "No"))]
317
+ crit = q.get("criteria", q.get("options"))
318
+ if qtype == "score":
319
+ opts = [str(c) for c in crit]
320
+ return qtype, text, list(range(len(opts))), opts
321
+ if isinstance(crit, dict):
322
+ return qtype, text, list(crit), [f"{k}: {v}" if v else str(k) for k, v in crit.items()]
323
+ return qtype, text, list(crit), [str(c) for c in crit]
324
+
325
+
326
+ @torch.no_grad()
327
+ def decide(model, tok, state, questions, defer_threshold=None, batch_size=32):
328
+ """Answer typed questions about one state.
329
+
330
+ questions: list of dicts, or a Jev/Laya-style dict {key: question}. Each question:
331
+ {"type": "choice"|"noul"|"score", "question"/"instructions": str,
332
+ "options": [...] or "criteria": {key: description} / [levels], "option_tokens"?: int}
333
+ """
334
+ if isinstance(questions, dict):
335
+ questions = [dict(q, key=k) for k, q in questions.items()]
336
+ items, metas = [], []
337
+ for q in questions:
338
+ qtype, text, keys, opts = _normalize_question(q)
339
+ item = {"state": state, "question": text, "options": opts, "type": qtype}
340
+ for extra in ("option_tokens", "seq_len"):
341
+ if extra in q:
342
+ item[extra] = q[extra]
343
+ items.append(item)
344
+ metas.append((q, qtype, keys, opts))
345
+ probs = score_items(model, tok, items, batch_size)
346
+ thr = model.fcfg.get("defer_threshold", 0.0) if defer_threshold is None else defer_threshold
347
+ results = []
348
+ for (q, qtype, keys, opts), p in zip(metas, probs):
349
+ i = int(np.argmax(p))
350
+ r = {"key": q.get("key"), "type": qtype, "question": items[len(results)]["question"],
351
+ "choice": keys[i], "choice_text": opts[i], "confidence": float(p[i]),
352
+ "probs": {str(k): float(v) for k, v in zip(keys, p)}, "defer": float(p[i]) < thr}
353
+ if qtype == "noul":
354
+ r["p_true"] = float(p[0])
355
+ if qtype == "score":
356
+ r["expected_level"] = float(np.dot(p, np.arange(len(p))))
357
+ results.append(r)
358
+ return {"results": results, "answers": {r["key"]: r for r in results if r["key"] is not None}}
falcondec_report.json ADDED
@@ -0,0 +1,1802 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "version": "1.0.0",
3
+ "notebook_version": "V3.7",
4
+ "mode": "scratch",
5
+ "parent_version": null,
6
+ "lineage": [
7
+ "jhu-clsp/ettin-encoder-150m"
8
+ ],
9
+ "config": {
10
+ "MODE": "scratch",
11
+ "FINETUNE_FROM": "latest",
12
+ "MODEL_NAME": "LightDec_V2_Long",
13
+ "BACKBONE": "jhu-clsp/ettin-encoder-150m",
14
+ "BACKBONE_INIT": "pretrained",
15
+ "PRESET": "long",
16
+ "CUSTOM_DATA_JSONL": "",
17
+ "CUSTOM_WEIGHT": 3.0,
18
+ "INCLUDE_BUILDERS": [],
19
+ "EXCLUDE_BUILDERS": [
20
+ "mind2web"
21
+ ],
22
+ "TEV1_DIR": "./tev1",
23
+ "TEV1_AUTOBUILD": true,
24
+ "TEV1_POLICY_CAP": 13500,
25
+ "TEV1_ROUTING_CAP": 6000,
26
+ "TEV1_RESEARCH_CAP": 1000,
27
+ "TRIAGE_EMAILS": 200,
28
+ "TEV1_RETRIES": 3,
29
+ "TEV1_STEP_TIMEOUT_MIN": 180,
30
+ "TEV1_STALL_MIN": 15,
31
+ "RETRY_FAILED_BUILDERS": true,
32
+ "LONG_TASK_TRAIN": 1500,
33
+ "MAX_LEN": 2048,
34
+ "LONG_MAX_LEN": 2048,
35
+ "LONG_OPTS_THRESHOLD": 24,
36
+ "HEAD_MAX_LEN": 192,
37
+ "MAX_TOK_PER_OPT": 24,
38
+ "MAX_OPTS_SINGLE_PASS": 96,
39
+ "INTERACT_LAYERS": 2,
40
+ "INTERACT_HEADS": 8,
41
+ "BATCH_SIZE": 128,
42
+ "TOKENS_PER_BATCH": 65536,
43
+ "GRAD_ACCUM": 1,
44
+ "LR_ENCODER": 8e-05,
45
+ "LR_HEAD": 0.0006,
46
+ "LLRD": 0.9,
47
+ "WEIGHT_DECAY": 0.01,
48
+ "WARMUP_FRAC": 0.06,
49
+ "EPOCHS": 0,
50
+ "EMA_DECAY": 0.999,
51
+ "GRAD_CLIP": 1.0,
52
+ "FINETUNE_LR_SCALE": 0.5,
53
+ "SPHERICAL_W": 0.5,
54
+ "BRIER_W": 0.0,
55
+ "RPS_W": 1.0,
56
+ "TASK_SAMPLING_ALPHA": 0.5,
57
+ "NOTA_PROB": 0.08,
58
+ "RLCD_EPOCHS": 0.0,
59
+ "RLCD_SAMPLES": 4,
60
+ "RLCD_SIGMA": 0.3,
61
+ "RLCD_CE_ANCHOR": 0.5,
62
+ "DEVICE": "cuda",
63
+ "SEED": 42,
64
+ "NUM_WORKERS": 8,
65
+ "USE_COMPILE": true,
66
+ "ATTN_IMPL": "auto",
67
+ "OUTPUT_ROOT": "./falcondec_runs",
68
+ "EXPORT_INT8": true,
69
+ "SIZE_TARGET_MB": 250,
70
+ "DEFER_THRESHOLD": 0.7,
71
+ "EVAL_BASELINE_PROOF_V2": true,
72
+ "BASELINE_MAX_PER_TASK": 120,
73
+ "RUN_CPU_BENCH": true,
74
+ "PUSH_TO_HUB": false,
75
+ "HUB_REPO": ""
76
+ },
77
+ "epochs": 10,
78
+ "best_epoch": 4,
79
+ "rlcd_kept": false,
80
+ "history": [
81
+ {
82
+ "epoch": 1,
83
+ "train_loss": 0.7058714814313827,
84
+ "train_acc": 0.7342094450703411,
85
+ "val_macro": 0.8188568656508199,
86
+ "val_micro": 0.820702479338843,
87
+ "val_nll": 0.4076740937270036
88
+ },
89
+ {
90
+ "epoch": 2,
91
+ "train_loss": 0.431469622218984,
92
+ "train_acc": 0.8473141073247583,
93
+ "val_macro": 0.8450872188411825,
94
+ "val_micro": 0.8444628099173553,
95
+ "val_nll": 0.3623910443100994
96
+ },
97
+ {
98
+ "epoch": 3,
99
+ "train_loss": 0.34276796075090776,
100
+ "train_acc": 0.8791940521680108,
101
+ "val_macro": 0.8501904675940947,
102
+ "val_micro": 0.8499173553719008,
103
+ "val_nll": 0.39227130460368825
104
+ },
105
+ {
106
+ "epoch": 4,
107
+ "train_loss": 0.2773282952523419,
108
+ "train_acc": 0.9039863196679513,
109
+ "val_macro": 0.8519172452047367,
110
+ "val_micro": 0.8519834710743802,
111
+ "val_nll": 0.4529550437184426
112
+ },
113
+ {
114
+ "epoch": 5,
115
+ "train_loss": 0.22030453415052323,
116
+ "train_acc": 0.925851485588573,
117
+ "val_macro": 0.8488092001804082,
118
+ "val_micro": 0.85,
119
+ "val_nll": 0.5369876145373608
120
+ },
121
+ {
122
+ "epoch": 6,
123
+ "train_loss": 0.16953414903032543,
124
+ "train_acc": 0.9449156076226879,
125
+ "val_macro": 0.8474279008050457,
126
+ "val_micro": 0.8488842975206612,
127
+ "val_nll": 0.6705887193589142
128
+ },
129
+ {
130
+ "epoch": 7,
131
+ "train_loss": 0.126976966611751,
132
+ "train_acc": 0.9602833527848939,
133
+ "val_macro": 0.8453322694829587,
134
+ "val_micro": 0.8476446280991735,
135
+ "val_nll": 0.8430423012207805
136
+ },
137
+ {
138
+ "epoch": 8,
139
+ "train_loss": 0.09470793687766035,
140
+ "train_acc": 0.9714013975720073,
141
+ "val_macro": 0.8451769009269958,
142
+ "val_micro": 0.8467355371900827,
143
+ "val_nll": 1.0336680755654943
144
+ },
145
+ {
146
+ "epoch": 9,
147
+ "train_loss": 0.0710737590306488,
148
+ "train_acc": 0.9794225926173407,
149
+ "val_macro": 0.8435159403610613,
150
+ "val_micro": 0.8458677685950413,
151
+ "val_nll": 1.210751206553627
152
+ },
153
+ {
154
+ "epoch": 10,
155
+ "train_loss": 0.05665995772408678,
156
+ "train_acc": 0.9843832093140154,
157
+ "val_macro": 0.844272704258198,
158
+ "val_micro": 0.8474380165289256,
159
+ "val_nll": 1.41872131482475
160
+ }
161
+ ],
162
+ "train_minutes": 291.67019440333047,
163
+ "temperatures": [
164
+ [
165
+ 2.37591814994812,
166
+ 2.0426506996154785,
167
+ 1.6455482244491577,
168
+ 1.8417285680770874
169
+ ],
170
+ [
171
+ 2.4885058403015137,
172
+ 2.4885058403015137,
173
+ 2.4885058403015137,
174
+ 2.4885058403015137
175
+ ],
176
+ [
177
+ 2.2925639152526855,
178
+ 2.2925639152526855,
179
+ 2.2925639152526855,
180
+ 2.2925639152526855
181
+ ]
182
+ ],
183
+ "data_counts": {
184
+ "train": 1023814,
185
+ "validation": 24200,
186
+ "test": 31990
187
+ },
188
+ "skipped_builders": [],
189
+ "test_overall": {
190
+ "n": 31990.0,
191
+ "micro_acc": 0.784057517974367,
192
+ "macro_acc": 0.7887006006875437,
193
+ "nll": 0.5630830966790886,
194
+ "brier": 0.2954194093899411,
195
+ "ece": 0.02856345933167103,
196
+ "aurc": 0.06778106632931055,
197
+ "score_mae": 0.5060575584353655,
198
+ "coverage@0.7": 0.6797436698968428,
199
+ "acc_on_covered@0.7": 0.9059094044607956
200
+ },
201
+ "test_per_domain": {
202
+ "agentic": 0.821117397177404,
203
+ "classification": 0.6115,
204
+ "code": 0.9108344674425007,
205
+ "guardrails": 0.7984073149310548,
206
+ "intents": 0.8505,
207
+ "long_context": 1.0,
208
+ "policy": 0.8061818181818182,
209
+ "reasoning": 0.7180877154615182,
210
+ "support": 0.9815181518151815,
211
+ "tev1_benchmark": 0.6541666666666667,
212
+ "workflows": 0.773
213
+ },
214
+ "test_per_task": [
215
+ {
216
+ "task": "ag_news/topic",
217
+ "n": 500,
218
+ "acc": 0.9,
219
+ "chance": 0.25,
220
+ "nll": 0.2674782864307199,
221
+ "brier": 0.14296993909824685,
222
+ "heldout": false,
223
+ "domain": "classification",
224
+ "ece": 0.05165432274341582,
225
+ "proof_v2 (card)": NaN,
226
+ "\u0394 vs proof_v2": NaN
227
+ },
228
+ {
229
+ "task": "agentharm/refuse",
230
+ "n": 416,
231
+ "acc": 0.6995192307692307,
232
+ "chance": 0.5,
233
+ "nll": 0.6008703433729422,
234
+ "brier": 0.4076762512706606,
235
+ "heldout": true,
236
+ "domain": "guardrails",
237
+ "ece": 0.08916855761064935,
238
+ "proof_v2 (card)": NaN,
239
+ "\u0394 vs proof_v2": NaN
240
+ },
241
+ {
242
+ "task": "agenttrek/finish_now",
243
+ "n": 151,
244
+ "acc": 0.7880794701986755,
245
+ "chance": 0.5,
246
+ "nll": 0.445978463758558,
247
+ "brier": 0.30066953632150406,
248
+ "heldout": false,
249
+ "domain": "agentic",
250
+ "ece": 0.07512128787324918,
251
+ "proof_v2 (card)": NaN,
252
+ "\u0394 vs proof_v2": NaN
253
+ },
254
+ {
255
+ "task": "agenttrek/next_action_type",
256
+ "n": 487,
257
+ "acc": 0.864476386036961,
258
+ "chance": 0.19054952576513148,
259
+ "nll": 0.3929848723152591,
260
+ "brier": 0.2184262054212701,
261
+ "heldout": false,
262
+ "domain": "agentic",
263
+ "ece": 0.09202881025827395,
264
+ "proof_v2 (card)": NaN,
265
+ "\u0394 vs proof_v2": NaN
266
+ },
267
+ {
268
+ "task": "anli/nli",
269
+ "n": 498,
270
+ "acc": 0.4859437751004016,
271
+ "chance": 0.33333333333333326,
272
+ "nll": 1.069221544292677,
273
+ "brier": 0.6544603492386665,
274
+ "heldout": false,
275
+ "domain": "reasoning",
276
+ "ece": 0.14589204983299517,
277
+ "proof_v2 (card)": NaN,
278
+ "\u0394 vs proof_v2": NaN
279
+ },
280
+ {
281
+ "task": "aqua_rat/math",
282
+ "n": 247,
283
+ "acc": 0.340080971659919,
284
+ "chance": 0.2,
285
+ "nll": 1.54525283087603,
286
+ "brier": 0.7739737508966391,
287
+ "heldout": false,
288
+ "domain": "reasoning",
289
+ "ece": 0.080350092668765,
290
+ "proof_v2 (card)": NaN,
291
+ "\u0394 vs proof_v2": NaN
292
+ },
293
+ {
294
+ "task": "arc_challenge/mcq",
295
+ "n": 500,
296
+ "acc": 0.49,
297
+ "chance": 0.25006666666666666,
298
+ "nll": 1.3021090898208785,
299
+ "brier": 0.6752390805252713,
300
+ "heldout": true,
301
+ "domain": "reasoning",
302
+ "ece": 0.1680490539073944,
303
+ "proof_v2 (card)": 0.308,
304
+ "\u0394 vs proof_v2": 0.182
305
+ },
306
+ {
307
+ "task": "arc_easy/mcq",
308
+ "n": 500,
309
+ "acc": 0.612,
310
+ "chance": 0.24980000000000002,
311
+ "nll": 0.9830388274090365,
312
+ "brier": 0.5264415342806388,
313
+ "heldout": true,
314
+ "domain": "reasoning",
315
+ "ece": 0.1215095258355141,
316
+ "proof_v2 (card)": 0.425,
317
+ "\u0394 vs proof_v2": 0.187
318
+ },
319
+ {
320
+ "task": "banking77/intent",
321
+ "n": 500,
322
+ "acc": 0.922,
323
+ "chance": 0.3214666666666666,
324
+ "nll": 0.23128961177584006,
325
+ "brier": 0.11756106243365176,
326
+ "heldout": true,
327
+ "domain": "intents",
328
+ "ece": 0.037492330908775316,
329
+ "proof_v2 (card)": 0.883,
330
+ "\u0394 vs proof_v2": 0.039000000000000035
331
+ },
332
+ {
333
+ "task": "banking77/intent_77",
334
+ "n": 300,
335
+ "acc": 0.58,
336
+ "chance": 0.012987012987012986,
337
+ "nll": 2.159930980304489,
338
+ "brier": 0.6431857202493357,
339
+ "heldout": true,
340
+ "domain": "intents",
341
+ "ece": 0.24079640249411272,
342
+ "proof_v2 (card)": NaN,
343
+ "\u0394 vs proof_v2": NaN
344
+ },
345
+ {
346
+ "task": "bigclonebench/clone",
347
+ "n": 500,
348
+ "acc": 0.962,
349
+ "chance": 0.5,
350
+ "nll": 0.1312969167594565,
351
+ "brier": 0.06444151702850223,
352
+ "heldout": false,
353
+ "domain": "code",
354
+ "ece": 0.03135248208045959,
355
+ "proof_v2 (card)": 0.383,
356
+ "\u0394 vs proof_v2": 0.579
357
+ },
358
+ {
359
+ "task": "bitext/category",
360
+ "n": 500,
361
+ "acc": 1.0,
362
+ "chance": 0.16128124098124094,
363
+ "nll": 0.0013673773490934594,
364
+ "brier": 0.0003828941821233322,
365
+ "heldout": false,
366
+ "domain": "support",
367
+ "ece": 0.001247308015823329,
368
+ "proof_v2 (card)": NaN,
369
+ "\u0394 vs proof_v2": NaN
370
+ },
371
+ {
372
+ "task": "bitext/route",
373
+ "n": 500,
374
+ "acc": 1.0,
375
+ "chance": 0.19016313131313134,
376
+ "nll": 0.00031508885690954,
377
+ "brier": 4.903548307607296e-06,
378
+ "heldout": false,
379
+ "domain": "support",
380
+ "ece": 0.0003137041330337764,
381
+ "proof_v2 (card)": 0.958,
382
+ "\u0394 vs proof_v2": 0.04200000000000004
383
+ },
384
+ {
385
+ "task": "boolq/yes_no",
386
+ "n": 500,
387
+ "acc": 0.822,
388
+ "chance": 0.5,
389
+ "nll": 0.4842173567947466,
390
+ "brier": 0.2840138305161832,
391
+ "heldout": false,
392
+ "domain": "reasoning",
393
+ "ece": 0.0795972980260849,
394
+ "proof_v2 (card)": 0.717,
395
+ "\u0394 vs proof_v2": 0.10499999999999998
396
+ },
397
+ {
398
+ "task": "civil_comments/toxic",
399
+ "n": 500,
400
+ "acc": 0.926,
401
+ "chance": 0.5,
402
+ "nll": 0.2234244500696659,
403
+ "brier": 0.1209681824839501,
404
+ "heldout": false,
405
+ "domain": "guardrails",
406
+ "ece": 0.05123518884181974,
407
+ "proof_v2 (card)": NaN,
408
+ "\u0394 vs proof_v2": NaN
409
+ },
410
+ {
411
+ "task": "clinc150/intent",
412
+ "n": 500,
413
+ "acc": 0.97,
414
+ "chance": 0.12404474984829318,
415
+ "nll": 0.09906445816022005,
416
+ "brier": 0.04996487012600131,
417
+ "heldout": false,
418
+ "domain": "intents",
419
+ "ece": 0.014944501101970714,
420
+ "proof_v2 (card)": 0.85,
421
+ "\u0394 vs proof_v2": 0.12
422
+ },
423
+ {
424
+ "task": "codexglue/code_to_doc",
425
+ "n": 504,
426
+ "acc": 0.9900793650793651,
427
+ "chance": 0.26233465608465617,
428
+ "nll": 0.032474144255576765,
429
+ "brier": 0.016158895201785088,
430
+ "heldout": false,
431
+ "domain": "code",
432
+ "ece": 0.009654374349685009,
433
+ "proof_v2 (card)": 0.969,
434
+ "\u0394 vs proof_v2": 0.021079365079365142
435
+ },
436
+ {
437
+ "task": "codexglue/doc_to_code",
438
+ "n": 504,
439
+ "acc": 0.9761904761904762,
440
+ "chance": 0.2786044973544973,
441
+ "nll": 0.05970113779317181,
442
+ "brier": 0.03072688989396561,
443
+ "heldout": false,
444
+ "domain": "code",
445
+ "ece": 0.012547755761752008,
446
+ "proof_v2 (card)": 0.961,
447
+ "\u0394 vs proof_v2": 0.015190476190476199
448
+ },
449
+ {
450
+ "task": "codexglue/func_name",
451
+ "n": 467,
452
+ "acc": 0.9550321199143469,
453
+ "chance": 0.25499643112062814,
454
+ "nll": 0.14961532151979784,
455
+ "brier": 0.07053876912275295,
456
+ "heldout": false,
457
+ "domain": "code",
458
+ "ece": 0.03948740715388316,
459
+ "proof_v2 (card)": 0.901,
460
+ "\u0394 vs proof_v2": 0.054032119914346866
461
+ },
462
+ {
463
+ "task": "codexglue/lang_id",
464
+ "n": 504,
465
+ "acc": 1.0,
466
+ "chance": 0.2354166666666667,
467
+ "nll": 0.0014352430686462487,
468
+ "brier": 1.2864369183890018e-05,
469
+ "heldout": false,
470
+ "domain": "code",
471
+ "ece": 0.0014309174721203188,
472
+ "proof_v2 (card)": 0.997,
473
+ "\u0394 vs proof_v2": 0.0030000000000000027
474
+ },
475
+ {
476
+ "task": "commonsense_qa/mcq",
477
+ "n": 493,
478
+ "acc": 0.6430020283975659,
479
+ "chance": 0.19999999999999998,
480
+ "nll": 0.9523430099929484,
481
+ "brier": 0.49435087368161623,
482
+ "heldout": false,
483
+ "domain": "reasoning",
484
+ "ece": 0.14308575696442233,
485
+ "proof_v2 (card)": 0.442,
486
+ "\u0394 vs proof_v2": 0.20100202839756592
487
+ },
488
+ {
489
+ "task": "counsel/critique_quality",
490
+ "n": 201,
491
+ "acc": 0.6268656716417911,
492
+ "chance": 0.3333333333333333,
493
+ "nll": 1.5466346353110312,
494
+ "brier": 0.6427543071457474,
495
+ "heldout": false,
496
+ "domain": "agentic",
497
+ "ece": 0.27062572442477023,
498
+ "proof_v2 (card)": NaN,
499
+ "\u0394 vs proof_v2": NaN
500
+ },
501
+ {
502
+ "task": "counsel/step_has_error",
503
+ "n": 201,
504
+ "acc": 0.8308457711442786,
505
+ "chance": 0.5,
506
+ "nll": 0.8236887436017318,
507
+ "brier": 0.3201374357369142,
508
+ "heldout": false,
509
+ "domain": "agentic",
510
+ "ece": 0.14827988011326954,
511
+ "proof_v2 (card)": NaN,
512
+ "\u0394 vs proof_v2": NaN
513
+ },
514
+ {
515
+ "task": "devign/vulnerability",
516
+ "n": 500,
517
+ "acc": 0.632,
518
+ "chance": 0.5,
519
+ "nll": 0.6074751503933221,
520
+ "brier": 0.42706214521965163,
521
+ "heldout": false,
522
+ "domain": "code",
523
+ "ece": 0.07018342161178594,
524
+ "proof_v2 (card)": 0.542,
525
+ "\u0394 vs proof_v2": 0.08999999999999997
526
+ },
527
+ {
528
+ "task": "emotion/6way",
529
+ "n": 500,
530
+ "acc": 0.472,
531
+ "chance": 0.16666666666666663,
532
+ "nll": 1.561316128242761,
533
+ "brier": 0.7356062387487318,
534
+ "heldout": true,
535
+ "domain": "classification",
536
+ "ece": 0.21054179659485817,
537
+ "proof_v2 (card)": NaN,
538
+ "\u0394 vs proof_v2": NaN
539
+ },
540
+ {
541
+ "task": "gsm8k/math",
542
+ "n": 500,
543
+ "acc": 0.7,
544
+ "chance": 0.25,
545
+ "nll": 0.6923719964642078,
546
+ "brier": 0.4014884319624292,
547
+ "heldout": false,
548
+ "domain": "reasoning",
549
+ "ece": 0.08541783547401426,
550
+ "proof_v2 (card)": 0.275,
551
+ "\u0394 vs proof_v2": 0.42499999999999993
552
+ },
553
+ {
554
+ "task": "hellaswag/continuation",
555
+ "n": 500,
556
+ "acc": 0.576,
557
+ "chance": 0.25,
558
+ "nll": 0.968790470642969,
559
+ "brier": 0.5288133586177948,
560
+ "heldout": false,
561
+ "domain": "reasoning",
562
+ "ece": 0.07543198531866073,
563
+ "proof_v2 (card)": NaN,
564
+ "\u0394 vs proof_v2": NaN
565
+ },
566
+ {
567
+ "task": "hotpotqa/comparison_yes_no",
568
+ "n": 26,
569
+ "acc": 0.9230769230769231,
570
+ "chance": 0.5,
571
+ "nll": 0.3400710394176153,
572
+ "brier": 0.15484679762361436,
573
+ "heldout": false,
574
+ "domain": "agentic",
575
+ "ece": 0.08232885369887719,
576
+ "proof_v2 (card)": NaN,
577
+ "\u0394 vs proof_v2": NaN
578
+ },
579
+ {
580
+ "task": "hotpotqa/retrieve",
581
+ "n": 497,
582
+ "acc": 0.8933601609657947,
583
+ "chance": 0.16741720162243304,
584
+ "nll": 0.3457089332123268,
585
+ "brier": 0.16216258689991922,
586
+ "heldout": false,
587
+ "domain": "agentic",
588
+ "ece": 0.04268557528854612,
589
+ "proof_v2 (card)": NaN,
590
+ "\u0394 vs proof_v2": NaN
591
+ },
592
+ {
593
+ "task": "humaneval/completion",
594
+ "n": 119,
595
+ "acc": 0.8487394957983193,
596
+ "chance": 0.39355742296918783,
597
+ "nll": 0.46361927195851294,
598
+ "brier": 0.2730045795813753,
599
+ "heldout": true,
600
+ "domain": "code",
601
+ "ece": 0.13989568708323633,
602
+ "proof_v2 (card)": 0.575,
603
+ "\u0394 vs proof_v2": 0.27373949579831935
604
+ },
605
+ {
606
+ "task": "jailbreak/detect",
607
+ "n": 262,
608
+ "acc": 0.9732824427480916,
609
+ "chance": 0.5,
610
+ "nll": 0.08469655378148465,
611
+ "brier": 0.044003931208438284,
612
+ "heldout": false,
613
+ "domain": "guardrails",
614
+ "ece": 0.030764950368240628,
615
+ "proof_v2 (card)": NaN,
616
+ "\u0394 vs proof_v2": NaN
617
+ },
618
+ {
619
+ "task": "long/contract_clause",
620
+ "n": 150,
621
+ "acc": 1.0,
622
+ "chance": 0.19999999999999996,
623
+ "nll": 0.0002789537957808837,
624
+ "brier": 4.274148862877054e-07,
625
+ "heldout": false,
626
+ "domain": "long_context",
627
+ "ece": 0.00027882695198055973,
628
+ "proof_v2 (card)": NaN,
629
+ "\u0394 vs proof_v2": NaN
630
+ },
631
+ {
632
+ "task": "long/email_thread",
633
+ "n": 150,
634
+ "acc": 1.0,
635
+ "chance": 0.25,
636
+ "nll": 0.0003266528345905802,
637
+ "brier": 2.717841913247229e-06,
638
+ "heldout": false,
639
+ "domain": "long_context",
640
+ "ece": 0.0003259424368540209,
641
+ "proof_v2 (card)": NaN,
642
+ "\u0394 vs proof_v2": NaN
643
+ },
644
+ {
645
+ "task": "long/service_log",
646
+ "n": 150,
647
+ "acc": 1.0,
648
+ "chance": 0.19999999999999996,
649
+ "nll": 0.0010035930563268873,
650
+ "brier": 2.0987077842670875e-06,
651
+ "heldout": false,
652
+ "domain": "long_context",
653
+ "ece": 0.0010028688112894146,
654
+ "proof_v2 (card)": NaN,
655
+ "\u0394 vs proof_v2": NaN
656
+ },
657
+ {
658
+ "task": "massive_en/intent",
659
+ "n": 500,
660
+ "acc": 0.93,
661
+ "chance": 0.12958727789085603,
662
+ "nll": 0.20636169074449573,
663
+ "brier": 0.10375277943487987,
664
+ "heldout": false,
665
+ "domain": "intents",
666
+ "ece": 0.0284088468849659,
667
+ "proof_v2 (card)": NaN,
668
+ "\u0394 vs proof_v2": NaN
669
+ },
670
+ {
671
+ "task": "mbpp/bugspot",
672
+ "n": 256,
673
+ "acc": 0.85546875,
674
+ "chance": 0.40559895833333326,
675
+ "nll": 0.3667814989712497,
676
+ "brier": 0.21843909652329055,
677
+ "heldout": false,
678
+ "domain": "code",
679
+ "ece": 0.037241218378767364,
680
+ "proof_v2 (card)": 0.475,
681
+ "\u0394 vs proof_v2": 0.38046875
682
+ },
683
+ {
684
+ "task": "mbpp/solution",
685
+ "n": 500,
686
+ "acc": 0.978,
687
+ "chance": 0.25,
688
+ "nll": 0.05863042902501184,
689
+ "brier": 0.02981698831494921,
690
+ "heldout": false,
691
+ "domain": "code",
692
+ "ece": 0.01960155099630357,
693
+ "proof_v2 (card)": 0.992,
694
+ "\u0394 vs proof_v2": -0.014000000000000012
695
+ },
696
+ {
697
+ "task": "mmlu/mcq",
698
+ "n": 500,
699
+ "acc": 0.39,
700
+ "chance": 0.25,
701
+ "nll": 1.4543649319559335,
702
+ "brier": 0.7702575193705765,
703
+ "heldout": true,
704
+ "domain": "reasoning",
705
+ "ece": 0.15507307499647138,
706
+ "proof_v2 (card)": NaN,
707
+ "\u0394 vs proof_v2": NaN
708
+ },
709
+ {
710
+ "task": "mnli/claim",
711
+ "n": 500,
712
+ "acc": 0.86,
713
+ "chance": 0.33333333333333326,
714
+ "nll": 0.4006310784481466,
715
+ "brier": 0.2131232579488494,
716
+ "heldout": false,
717
+ "domain": "reasoning",
718
+ "ece": 0.081218329668045,
719
+ "proof_v2 (card)": 0.492,
720
+ "\u0394 vs proof_v2": 0.368
721
+ },
722
+ {
723
+ "task": "openbookqa/mcq",
724
+ "n": 500,
725
+ "acc": 0.572,
726
+ "chance": 0.25,
727
+ "nll": 1.1920220527790952,
728
+ "brier": 0.6155284183368397,
729
+ "heldout": false,
730
+ "domain": "reasoning",
731
+ "ece": 0.21775564682483672,
732
+ "proof_v2 (card)": 0.292,
733
+ "\u0394 vs proof_v2": 0.27999999999999997
734
+ },
735
+ {
736
+ "task": "policy/access_control_transfer",
737
+ "n": 500,
738
+ "acc": 1.0,
739
+ "chance": 0.33333333333333326,
740
+ "nll": 0.002006467206090747,
741
+ "brier": 2.4728843418299015e-05,
742
+ "heldout": false,
743
+ "domain": "policy",
744
+ "ece": 0.001999327063560541,
745
+ "proof_v2 (card)": NaN,
746
+ "\u0394 vs proof_v2": NaN
747
+ },
748
+ {
749
+ "task": "policy/count_threshold_transfer",
750
+ "n": 500,
751
+ "acc": 0.88,
752
+ "chance": 0.10537052392052391,
753
+ "nll": 0.319376940273738,
754
+ "brier": 0.1809248790341789,
755
+ "heldout": false,
756
+ "domain": "policy",
757
+ "ece": 0.04758520478010175,
758
+ "proof_v2 (card)": NaN,
759
+ "\u0394 vs proof_v2": NaN
760
+ },
761
+ {
762
+ "task": "policy/free_shipping_transfer",
763
+ "n": 500,
764
+ "acc": 0.938,
765
+ "chance": 0.5,
766
+ "nll": 0.1287872482436942,
767
+ "brier": 0.08322479293591872,
768
+ "heldout": false,
769
+ "domain": "policy",
770
+ "ece": 0.041308605194091734,
771
+ "proof_v2 (card)": NaN,
772
+ "\u0394 vs proof_v2": NaN
773
+ },
774
+ {
775
+ "task": "policy/invoice_overdue_transfer",
776
+ "n": 500,
777
+ "acc": 0.824,
778
+ "chance": 0.5,
779
+ "nll": 0.3965394942490384,
780
+ "brier": 0.2504247577332596,
781
+ "heldout": false,
782
+ "domain": "policy",
783
+ "ece": 0.01985593855381013,
784
+ "proof_v2 (card)": NaN,
785
+ "\u0394 vs proof_v2": NaN
786
+ },
787
+ {
788
+ "task": "policy/invoice_total_transfer",
789
+ "n": 500,
790
+ "acc": 0.474,
791
+ "chance": 0.5,
792
+ "nll": 0.710681935429573,
793
+ "brier": 0.5164436840846652,
794
+ "heldout": false,
795
+ "domain": "policy",
796
+ "ece": 0.08196303224563599,
797
+ "proof_v2 (card)": NaN,
798
+ "\u0394 vs proof_v2": NaN
799
+ },
800
+ {
801
+ "task": "policy/refund_approval_transfer",
802
+ "n": 500,
803
+ "acc": 0.962,
804
+ "chance": 0.33333333333333326,
805
+ "nll": 0.10006511305032836,
806
+ "brier": 0.057969574647252664,
807
+ "heldout": false,
808
+ "domain": "policy",
809
+ "ece": 0.018834196925163266,
810
+ "proof_v2 (card)": NaN,
811
+ "\u0394 vs proof_v2": NaN
812
+ },
813
+ {
814
+ "task": "policy/return_window_transfer",
815
+ "n": 500,
816
+ "acc": 1.0,
817
+ "chance": 0.33333333333333326,
818
+ "nll": 0.02710788446944207,
819
+ "brier": 0.004185976644178241,
820
+ "heldout": false,
821
+ "domain": "policy",
822
+ "ece": 0.02567652928829193,
823
+ "proof_v2 (card)": NaN,
824
+ "\u0394 vs proof_v2": NaN
825
+ },
826
+ {
827
+ "task": "policy/sla_urgency_transfer",
828
+ "n": 500,
829
+ "acc": 0.606,
830
+ "chance": 0.25,
831
+ "nll": 0.8170870101451874,
832
+ "brier": 0.5123369210022903,
833
+ "heldout": false,
834
+ "domain": "policy",
835
+ "ece": 0.13068743020296097,
836
+ "proof_v2 (card)": NaN,
837
+ "\u0394 vs proof_v2": NaN
838
+ },
839
+ {
840
+ "task": "policy/table_compare_transfer",
841
+ "n": 500,
842
+ "acc": 0.928,
843
+ "chance": 0.5,
844
+ "nll": 0.2105378701629961,
845
+ "brier": 0.11898007306418568,
846
+ "heldout": false,
847
+ "domain": "policy",
848
+ "ece": 0.04929606854915622,
849
+ "proof_v2 (card)": NaN,
850
+ "\u0394 vs proof_v2": NaN
851
+ },
852
+ {
853
+ "task": "policy/table_count_transfer",
854
+ "n": 500,
855
+ "acc": 0.324,
856
+ "chance": 0.12114228549228548,
857
+ "nll": 3.24442003638536,
858
+ "brier": 1.1488026091975905,
859
+ "heldout": false,
860
+ "domain": "policy",
861
+ "ece": 0.5407210965156555,
862
+ "proof_v2 (card)": NaN,
863
+ "\u0394 vs proof_v2": NaN
864
+ },
865
+ {
866
+ "task": "policy/table_extreme_transfer",
867
+ "n": 500,
868
+ "acc": 0.932,
869
+ "chance": 0.12210815295815294,
870
+ "nll": 0.268062693382595,
871
+ "brier": 0.11322196827103552,
872
+ "heldout": false,
873
+ "domain": "policy",
874
+ "ece": 0.05259446609020234,
875
+ "proof_v2 (card)": NaN,
876
+ "\u0394 vs proof_v2": NaN
877
+ },
878
+ {
879
+ "task": "prompt_injections/detect",
880
+ "n": 116,
881
+ "acc": 0.5948275862068966,
882
+ "chance": 0.5,
883
+ "nll": 1.1838965486195179,
884
+ "brier": 0.6826645118148511,
885
+ "heldout": true,
886
+ "domain": "guardrails",
887
+ "ece": 0.33272201953263114,
888
+ "proof_v2 (card)": NaN,
889
+ "\u0394 vs proof_v2": NaN
890
+ },
891
+ {
892
+ "task": "qasc/mcq",
893
+ "n": 500,
894
+ "acc": 0.986,
895
+ "chance": 0.125,
896
+ "nll": 0.028506858559036514,
897
+ "brier": 0.017274368003708154,
898
+ "heldout": false,
899
+ "domain": "reasoning",
900
+ "ece": 0.007774531126022318,
901
+ "proof_v2 (card)": NaN,
902
+ "\u0394 vs proof_v2": NaN
903
+ },
904
+ {
905
+ "task": "sciq/mcq",
906
+ "n": 498,
907
+ "acc": 0.9538152610441767,
908
+ "chance": 0.25,
909
+ "nll": 0.1506263591345834,
910
+ "brier": 0.06985463059035983,
911
+ "heldout": false,
912
+ "domain": "reasoning",
913
+ "ece": 0.020590302994452258,
914
+ "proof_v2 (card)": 0.692,
915
+ "\u0394 vs proof_v2": 0.26181526104417674
916
+ },
917
+ {
918
+ "task": "scitail/support",
919
+ "n": 500,
920
+ "acc": 0.962,
921
+ "chance": 0.5,
922
+ "nll": 0.11612921302905306,
923
+ "brier": 0.05990850379879065,
924
+ "heldout": false,
925
+ "domain": "reasoning",
926
+ "ece": 0.03015338957309728,
927
+ "proof_v2 (card)": NaN,
928
+ "\u0394 vs proof_v2": NaN
929
+ },
930
+ {
931
+ "task": "snli/contradicts",
932
+ "n": 500,
933
+ "acc": 0.99,
934
+ "chance": 0.33333333333333326,
935
+ "nll": 0.042843449617153966,
936
+ "brier": 0.017027529190616314,
937
+ "heldout": false,
938
+ "domain": "reasoning",
939
+ "ece": 0.018511961221695038,
940
+ "proof_v2 (card)": 0.892,
941
+ "\u0394 vs proof_v2": 0.09799999999999998
942
+ },
943
+ {
944
+ "task": "snli/must_be_true",
945
+ "n": 500,
946
+ "acc": 0.986,
947
+ "chance": 0.33333333333333326,
948
+ "nll": 0.06753946821717545,
949
+ "brier": 0.02560326154020915,
950
+ "heldout": false,
951
+ "domain": "reasoning",
952
+ "ece": 0.043202445626258815,
953
+ "proof_v2 (card)": 0.908,
954
+ "\u0394 vs proof_v2": 0.07799999999999996
955
+ },
956
+ {
957
+ "task": "snli/nli",
958
+ "n": 988,
959
+ "acc": 0.8947368421052632,
960
+ "chance": 0.33333333333333326,
961
+ "nll": 0.3554420507991845,
962
+ "brier": 0.17963376213248355,
963
+ "heldout": false,
964
+ "domain": "reasoning",
965
+ "ece": 0.10791712172842222,
966
+ "proof_v2 (card)": NaN,
967
+ "\u0394 vs proof_v2": NaN
968
+ },
969
+ {
970
+ "task": "sst5/score",
971
+ "n": 500,
972
+ "acc": 0.406,
973
+ "chance": 0.2,
974
+ "nll": 1.3002369542717933,
975
+ "brier": 0.6888741553408958,
976
+ "heldout": true,
977
+ "domain": "classification",
978
+ "ece": 0.068851743131876,
979
+ "proof_v2 (card)": NaN,
980
+ "\u0394 vs proof_v2": NaN
981
+ },
982
+ {
983
+ "task": "tev1_test/ag_news",
984
+ "n": 150,
985
+ "acc": 0.92,
986
+ "chance": 0.25,
987
+ "nll": 0.2815313008365532,
988
+ "brier": 0.13456139206846243,
989
+ "heldout": true,
990
+ "domain": "tev1_benchmark",
991
+ "ece": 0.09770219445228576,
992
+ "proof_v2 (card)": NaN,
993
+ "\u0394 vs proof_v2": NaN
994
+ },
995
+ {
996
+ "task": "tev1_test/banking77",
997
+ "n": 200,
998
+ "acc": 0.72,
999
+ "chance": 0.17791666666666664,
1000
+ "nll": 0.777236735031438,
1001
+ "brier": 0.3902663054271286,
1002
+ "heldout": true,
1003
+ "domain": "tev1_benchmark",
1004
+ "ece": 0.13721376925706863,
1005
+ "proof_v2 (card)": NaN,
1006
+ "\u0394 vs proof_v2": NaN
1007
+ },
1008
+ {
1009
+ "task": "tev1_test/boolq",
1010
+ "n": 200,
1011
+ "acc": 0.835,
1012
+ "chance": 0.5,
1013
+ "nll": 0.40902234488283284,
1014
+ "brier": 0.24753055362669119,
1015
+ "heldout": true,
1016
+ "domain": "tev1_benchmark",
1017
+ "ece": 0.04955918818712238,
1018
+ "proof_v2 (card)": NaN,
1019
+ "\u0394 vs proof_v2": NaN
1020
+ },
1021
+ {
1022
+ "task": "tev1_test/mnli",
1023
+ "n": 300,
1024
+ "acc": 0.7233333333333334,
1025
+ "chance": 0.3333333333333333,
1026
+ "nll": 0.6736717029288412,
1027
+ "brier": 0.39782655881245704,
1028
+ "heldout": true,
1029
+ "domain": "tev1_benchmark",
1030
+ "ece": 0.07533363938331603,
1031
+ "proof_v2 (card)": NaN,
1032
+ "\u0394 vs proof_v2": NaN
1033
+ },
1034
+ {
1035
+ "task": "tev1_test/policy",
1036
+ "n": 1200,
1037
+ "acc": 0.5291666666666667,
1038
+ "chance": 0.3333333333333333,
1039
+ "nll": 0.8761954319352905,
1040
+ "brier": 0.5462664224307302,
1041
+ "heldout": true,
1042
+ "domain": "tev1_benchmark",
1043
+ "ece": 0.05868758827447889,
1044
+ "proof_v2 (card)": NaN,
1045
+ "\u0394 vs proof_v2": NaN
1046
+ },
1047
+ {
1048
+ "task": "tev1_test/routing",
1049
+ "n": 600,
1050
+ "acc": 0.4583333333333333,
1051
+ "chance": 0.2,
1052
+ "nll": 1.120521725914441,
1053
+ "brier": 0.5789269511485964,
1054
+ "heldout": true,
1055
+ "domain": "tev1_benchmark",
1056
+ "ece": 0.037736288358767814,
1057
+ "proof_v2 (card)": NaN,
1058
+ "\u0394 vs proof_v2": NaN
1059
+ },
1060
+ {
1061
+ "task": "tev1_test/sst5",
1062
+ "n": 150,
1063
+ "acc": 0.3933333333333333,
1064
+ "chance": 0.19999999999999996,
1065
+ "nll": 1.294863009850184,
1066
+ "brier": 0.6810328679494898,
1067
+ "heldout": true,
1068
+ "domain": "tev1_benchmark",
1069
+ "ece": 0.1081918982664744,
1070
+ "proof_v2 (card)": NaN,
1071
+ "\u0394 vs proof_v2": NaN
1072
+ },
1073
+ {
1074
+ "task": "triage/support_email",
1075
+ "n": 505,
1076
+ "acc": 0.9445544554455445,
1077
+ "chance": 0.32333333333333325,
1078
+ "nll": 0.15705669974757994,
1079
+ "brier": 0.07848840114983771,
1080
+ "heldout": false,
1081
+ "domain": "support",
1082
+ "ece": 0.045869263741049465,
1083
+ "proof_v2 (card)": NaN,
1084
+ "\u0394 vs proof_v2": NaN
1085
+ },
1086
+ {
1087
+ "task": "typed_decisions/agent_trace_observability",
1088
+ "n": 500,
1089
+ "acc": 0.734,
1090
+ "chance": 0.3,
1091
+ "nll": 0.8089503426551818,
1092
+ "brier": 0.45125534421065666,
1093
+ "heldout": false,
1094
+ "domain": "workflows",
1095
+ "ece": 0.22166652178764346,
1096
+ "proof_v2 (card)": NaN,
1097
+ "\u0394 vs proof_v2": NaN
1098
+ },
1099
+ {
1100
+ "task": "typed_decisions/customer_service",
1101
+ "n": 500,
1102
+ "acc": 0.764,
1103
+ "chance": 0.28,
1104
+ "nll": 0.7428050636351109,
1105
+ "brier": 0.40177602517263195,
1106
+ "heldout": false,
1107
+ "domain": "workflows",
1108
+ "ece": 0.20972245779633528,
1109
+ "proof_v2 (card)": NaN,
1110
+ "\u0394 vs proof_v2": NaN
1111
+ },
1112
+ {
1113
+ "task": "typed_decisions/invoice_processing",
1114
+ "n": 500,
1115
+ "acc": 0.826,
1116
+ "chance": 0.35,
1117
+ "nll": 0.5732684296518564,
1118
+ "brier": 0.2997603036136731,
1119
+ "heldout": false,
1120
+ "domain": "workflows",
1121
+ "ece": 0.19463082921504973,
1122
+ "proof_v2 (card)": NaN,
1123
+ "\u0394 vs proof_v2": NaN
1124
+ },
1125
+ {
1126
+ "task": "typed_decisions/security_incidents",
1127
+ "n": 500,
1128
+ "acc": 0.768,
1129
+ "chance": 0.34,
1130
+ "nll": 0.7502468670606613,
1131
+ "brier": 0.4238190137148829,
1132
+ "heldout": false,
1133
+ "domain": "workflows",
1134
+ "ece": 0.23508695220947262,
1135
+ "proof_v2 (card)": NaN,
1136
+ "\u0394 vs proof_v2": NaN
1137
+ },
1138
+ {
1139
+ "task": "winogrande/blank",
1140
+ "n": 500,
1141
+ "acc": 0.662,
1142
+ "chance": 0.5,
1143
+ "nll": 0.6651458515003323,
1144
+ "brier": 0.45323706188140617,
1145
+ "heldout": false,
1146
+ "domain": "reasoning",
1147
+ "ece": 0.1159557296037674,
1148
+ "proof_v2 (card)": NaN,
1149
+ "\u0394 vs proof_v2": NaN
1150
+ },
1151
+ {
1152
+ "task": "yelp/score",
1153
+ "n": 500,
1154
+ "acc": 0.668,
1155
+ "chance": 0.2,
1156
+ "nll": 0.7834262755662202,
1157
+ "brier": 0.4423740240858843,
1158
+ "heldout": false,
1159
+ "domain": "classification",
1160
+ "ece": 0.09068749487400055,
1161
+ "proof_v2 (card)": NaN,
1162
+ "\u0394 vs proof_v2": NaN
1163
+ }
1164
+ ],
1165
+ "head_to_head": [
1166
+ {
1167
+ "task": "ag_news/topic",
1168
+ "n": 120,
1169
+ "LightDec": 0.85,
1170
+ "proof_v2": 0.45,
1171
+ "\u0394": 0.39999999999999997
1172
+ },
1173
+ {
1174
+ "task": "agentharm/refuse",
1175
+ "n": 120,
1176
+ "LightDec": 0.6666666666666666,
1177
+ "proof_v2": 0.45,
1178
+ "\u0394": 0.21666666666666662
1179
+ },
1180
+ {
1181
+ "task": "agenttrek/finish_now",
1182
+ "n": 120,
1183
+ "LightDec": 0.7833333333333333,
1184
+ "proof_v2": 0.25833333333333336,
1185
+ "\u0394": 0.5249999999999999
1186
+ },
1187
+ {
1188
+ "task": "agenttrek/next_action_type",
1189
+ "n": 120,
1190
+ "LightDec": 0.875,
1191
+ "proof_v2": 0.075,
1192
+ "\u0394": 0.8
1193
+ },
1194
+ {
1195
+ "task": "anli/nli",
1196
+ "n": 120,
1197
+ "LightDec": 0.5916666666666667,
1198
+ "proof_v2": 0.325,
1199
+ "\u0394": 0.26666666666666666
1200
+ },
1201
+ {
1202
+ "task": "aqua_rat/math",
1203
+ "n": 120,
1204
+ "LightDec": 0.35,
1205
+ "proof_v2": 0.24166666666666667,
1206
+ "\u0394": 0.10833333333333331
1207
+ },
1208
+ {
1209
+ "task": "arc_challenge/mcq",
1210
+ "n": 120,
1211
+ "LightDec": 0.5333333333333333,
1212
+ "proof_v2": 0.30833333333333335,
1213
+ "\u0394": 0.22499999999999998
1214
+ },
1215
+ {
1216
+ "task": "arc_easy/mcq",
1217
+ "n": 120,
1218
+ "LightDec": 0.65,
1219
+ "proof_v2": 0.4583333333333333,
1220
+ "\u0394": 0.1916666666666667
1221
+ },
1222
+ {
1223
+ "task": "banking77/intent",
1224
+ "n": 120,
1225
+ "LightDec": 0.925,
1226
+ "proof_v2": 0.8666666666666667,
1227
+ "\u0394": 0.05833333333333335
1228
+ },
1229
+ {
1230
+ "task": "bigclonebench/clone",
1231
+ "n": 120,
1232
+ "LightDec": 0.9666666666666667,
1233
+ "proof_v2": 0.24166666666666667,
1234
+ "\u0394": 0.725
1235
+ },
1236
+ {
1237
+ "task": "bitext/category",
1238
+ "n": 120,
1239
+ "LightDec": 1.0,
1240
+ "proof_v2": 0.8583333333333333,
1241
+ "\u0394": 0.14166666666666672
1242
+ },
1243
+ {
1244
+ "task": "bitext/route",
1245
+ "n": 120,
1246
+ "LightDec": 1.0,
1247
+ "proof_v2": 0.925,
1248
+ "\u0394": 0.07499999999999996
1249
+ },
1250
+ {
1251
+ "task": "boolq/yes_no",
1252
+ "n": 120,
1253
+ "LightDec": 0.825,
1254
+ "proof_v2": 0.6666666666666666,
1255
+ "\u0394": 0.15833333333333333
1256
+ },
1257
+ {
1258
+ "task": "civil_comments/toxic",
1259
+ "n": 120,
1260
+ "LightDec": 0.9333333333333333,
1261
+ "proof_v2": 0.35833333333333334,
1262
+ "\u0394": 0.575
1263
+ },
1264
+ {
1265
+ "task": "clinc150/intent",
1266
+ "n": 120,
1267
+ "LightDec": 0.975,
1268
+ "proof_v2": 0.6416666666666667,
1269
+ "\u0394": 0.33333333333333326
1270
+ },
1271
+ {
1272
+ "task": "codexglue/code_to_doc",
1273
+ "n": 120,
1274
+ "LightDec": 1.0,
1275
+ "proof_v2": 0.9333333333333333,
1276
+ "\u0394": 0.06666666666666665
1277
+ },
1278
+ {
1279
+ "task": "codexglue/doc_to_code",
1280
+ "n": 120,
1281
+ "LightDec": 0.9833333333333333,
1282
+ "proof_v2": 0.975,
1283
+ "\u0394": 0.008333333333333304
1284
+ },
1285
+ {
1286
+ "task": "codexglue/func_name",
1287
+ "n": 120,
1288
+ "LightDec": 0.9916666666666667,
1289
+ "proof_v2": 0.9166666666666666,
1290
+ "\u0394": 0.07500000000000007
1291
+ },
1292
+ {
1293
+ "task": "codexglue/lang_id",
1294
+ "n": 120,
1295
+ "LightDec": 1.0,
1296
+ "proof_v2": 1.0,
1297
+ "\u0394": 0.0
1298
+ },
1299
+ {
1300
+ "task": "commonsense_qa/mcq",
1301
+ "n": 120,
1302
+ "LightDec": 0.6833333333333333,
1303
+ "proof_v2": 0.4083333333333333,
1304
+ "\u0394": 0.275
1305
+ },
1306
+ {
1307
+ "task": "counsel/critique_quality",
1308
+ "n": 120,
1309
+ "LightDec": 0.6083333333333333,
1310
+ "proof_v2": 0.25833333333333336,
1311
+ "\u0394": 0.3499999999999999
1312
+ },
1313
+ {
1314
+ "task": "counsel/step_has_error",
1315
+ "n": 120,
1316
+ "LightDec": 0.825,
1317
+ "proof_v2": 0.75,
1318
+ "\u0394": 0.07499999999999996
1319
+ },
1320
+ {
1321
+ "task": "devign/vulnerability",
1322
+ "n": 120,
1323
+ "LightDec": 0.6666666666666666,
1324
+ "proof_v2": 0.5666666666666667,
1325
+ "\u0394": 0.09999999999999998
1326
+ },
1327
+ {
1328
+ "task": "emotion/6way",
1329
+ "n": 120,
1330
+ "LightDec": 0.49166666666666664,
1331
+ "proof_v2": 0.36666666666666664,
1332
+ "\u0394": 0.125
1333
+ },
1334
+ {
1335
+ "task": "gsm8k/math",
1336
+ "n": 120,
1337
+ "LightDec": 0.7333333333333333,
1338
+ "proof_v2": 0.125,
1339
+ "\u0394": 0.6083333333333333
1340
+ },
1341
+ {
1342
+ "task": "hellaswag/continuation",
1343
+ "n": 120,
1344
+ "LightDec": 0.6083333333333333,
1345
+ "proof_v2": 0.2916666666666667,
1346
+ "\u0394": 0.3166666666666666
1347
+ },
1348
+ {
1349
+ "task": "hotpotqa/comparison_yes_no",
1350
+ "n": 26,
1351
+ "LightDec": 0.9230769230769231,
1352
+ "proof_v2": 0.46153846153846156,
1353
+ "\u0394": 0.46153846153846156
1354
+ },
1355
+ {
1356
+ "task": "hotpotqa/retrieve",
1357
+ "n": 120,
1358
+ "LightDec": 0.875,
1359
+ "proof_v2": 0.2916666666666667,
1360
+ "\u0394": 0.5833333333333333
1361
+ },
1362
+ {
1363
+ "task": "humaneval/completion",
1364
+ "n": 119,
1365
+ "LightDec": 0.8487394957983193,
1366
+ "proof_v2": 0.5126050420168067,
1367
+ "\u0394": 0.33613445378151263
1368
+ },
1369
+ {
1370
+ "task": "jailbreak/detect",
1371
+ "n": 120,
1372
+ "LightDec": 0.9833333333333333,
1373
+ "proof_v2": 0.5666666666666667,
1374
+ "\u0394": 0.41666666666666663
1375
+ },
1376
+ {
1377
+ "task": "long/contract_clause",
1378
+ "n": 120,
1379
+ "LightDec": 1.0,
1380
+ "proof_v2": 0.23333333333333334,
1381
+ "\u0394": 0.7666666666666666
1382
+ },
1383
+ {
1384
+ "task": "long/email_thread",
1385
+ "n": 120,
1386
+ "LightDec": 1.0,
1387
+ "proof_v2": 0.05,
1388
+ "\u0394": 0.95
1389
+ },
1390
+ {
1391
+ "task": "long/service_log",
1392
+ "n": 120,
1393
+ "LightDec": 1.0,
1394
+ "proof_v2": 0.24166666666666667,
1395
+ "\u0394": 0.7583333333333333
1396
+ },
1397
+ {
1398
+ "task": "massive_en/intent",
1399
+ "n": 120,
1400
+ "LightDec": 0.95,
1401
+ "proof_v2": 0.7416666666666667,
1402
+ "\u0394": 0.20833333333333326
1403
+ },
1404
+ {
1405
+ "task": "mbpp/bugspot",
1406
+ "n": 120,
1407
+ "LightDec": 0.8583333333333333,
1408
+ "proof_v2": 0.575,
1409
+ "\u0394": 0.2833333333333333
1410
+ },
1411
+ {
1412
+ "task": "mbpp/solution",
1413
+ "n": 120,
1414
+ "LightDec": 1.0,
1415
+ "proof_v2": 0.85,
1416
+ "\u0394": 0.15000000000000002
1417
+ },
1418
+ {
1419
+ "task": "mmlu/mcq",
1420
+ "n": 120,
1421
+ "LightDec": 0.4083333333333333,
1422
+ "proof_v2": 0.30833333333333335,
1423
+ "\u0394": 0.09999999999999998
1424
+ },
1425
+ {
1426
+ "task": "mnli/claim",
1427
+ "n": 120,
1428
+ "LightDec": 0.875,
1429
+ "proof_v2": 0.31666666666666665,
1430
+ "\u0394": 0.5583333333333333
1431
+ },
1432
+ {
1433
+ "task": "openbookqa/mcq",
1434
+ "n": 120,
1435
+ "LightDec": 0.65,
1436
+ "proof_v2": 0.2833333333333333,
1437
+ "\u0394": 0.3666666666666667
1438
+ },
1439
+ {
1440
+ "task": "policy/access_control_transfer",
1441
+ "n": 120,
1442
+ "LightDec": 1.0,
1443
+ "proof_v2": 0.23333333333333334,
1444
+ "\u0394": 0.7666666666666666
1445
+ },
1446
+ {
1447
+ "task": "policy/count_threshold_transfer",
1448
+ "n": 120,
1449
+ "LightDec": 0.8666666666666667,
1450
+ "proof_v2": 0.10833333333333334,
1451
+ "\u0394": 0.7583333333333333
1452
+ },
1453
+ {
1454
+ "task": "policy/free_shipping_transfer",
1455
+ "n": 120,
1456
+ "LightDec": 0.925,
1457
+ "proof_v2": 0.7083333333333334,
1458
+ "\u0394": 0.21666666666666667
1459
+ },
1460
+ {
1461
+ "task": "policy/invoice_overdue_transfer",
1462
+ "n": 120,
1463
+ "LightDec": 0.8083333333333333,
1464
+ "proof_v2": 0.6583333333333333,
1465
+ "\u0394": 0.15000000000000002
1466
+ },
1467
+ {
1468
+ "task": "policy/invoice_total_transfer",
1469
+ "n": 120,
1470
+ "LightDec": 0.45,
1471
+ "proof_v2": 0.525,
1472
+ "\u0394": -0.07500000000000001
1473
+ },
1474
+ {
1475
+ "task": "policy/refund_approval_transfer",
1476
+ "n": 120,
1477
+ "LightDec": 0.9583333333333334,
1478
+ "proof_v2": 0.325,
1479
+ "\u0394": 0.6333333333333333
1480
+ },
1481
+ {
1482
+ "task": "policy/return_window_transfer",
1483
+ "n": 120,
1484
+ "LightDec": 1.0,
1485
+ "proof_v2": 0.3416666666666667,
1486
+ "\u0394": 0.6583333333333333
1487
+ },
1488
+ {
1489
+ "task": "policy/sla_urgency_transfer",
1490
+ "n": 120,
1491
+ "LightDec": 0.5166666666666667,
1492
+ "proof_v2": 0.44166666666666665,
1493
+ "\u0394": 0.07500000000000007
1494
+ },
1495
+ {
1496
+ "task": "policy/table_compare_transfer",
1497
+ "n": 120,
1498
+ "LightDec": 0.9166666666666666,
1499
+ "proof_v2": 0.49166666666666664,
1500
+ "\u0394": 0.425
1501
+ },
1502
+ {
1503
+ "task": "policy/table_count_transfer",
1504
+ "n": 120,
1505
+ "LightDec": 0.31666666666666665,
1506
+ "proof_v2": 0.15833333333333333,
1507
+ "\u0394": 0.15833333333333333
1508
+ },
1509
+ {
1510
+ "task": "policy/table_extreme_transfer",
1511
+ "n": 120,
1512
+ "LightDec": 0.925,
1513
+ "proof_v2": 0.14166666666666666,
1514
+ "\u0394": 0.7833333333333334
1515
+ },
1516
+ {
1517
+ "task": "prompt_injections/detect",
1518
+ "n": 116,
1519
+ "LightDec": 0.5948275862068966,
1520
+ "proof_v2": 0.5431034482758621,
1521
+ "\u0394": 0.051724137931034475
1522
+ },
1523
+ {
1524
+ "task": "qasc/mcq",
1525
+ "n": 120,
1526
+ "LightDec": 0.9916666666666667,
1527
+ "proof_v2": 0.8583333333333333,
1528
+ "\u0394": 0.13333333333333341
1529
+ },
1530
+ {
1531
+ "task": "sciq/mcq",
1532
+ "n": 120,
1533
+ "LightDec": 0.9666666666666667,
1534
+ "proof_v2": 0.8666666666666667,
1535
+ "\u0394": 0.09999999999999998
1536
+ },
1537
+ {
1538
+ "task": "scitail/support",
1539
+ "n": 120,
1540
+ "LightDec": 0.9583333333333334,
1541
+ "proof_v2": 0.425,
1542
+ "\u0394": 0.5333333333333334
1543
+ },
1544
+ {
1545
+ "task": "snli/contradicts",
1546
+ "n": 120,
1547
+ "LightDec": 1.0,
1548
+ "proof_v2": 0.35833333333333334,
1549
+ "\u0394": 0.6416666666666666
1550
+ },
1551
+ {
1552
+ "task": "snli/must_be_true",
1553
+ "n": 120,
1554
+ "LightDec": 0.9916666666666667,
1555
+ "proof_v2": 0.925,
1556
+ "\u0394": 0.06666666666666665
1557
+ },
1558
+ {
1559
+ "task": "snli/nli",
1560
+ "n": 120,
1561
+ "LightDec": 0.9083333333333333,
1562
+ "proof_v2": 0.3416666666666667,
1563
+ "\u0394": 0.5666666666666667
1564
+ },
1565
+ {
1566
+ "task": "sst5/score",
1567
+ "n": 120,
1568
+ "LightDec": 0.38333333333333336,
1569
+ "proof_v2": 0.2833333333333333,
1570
+ "\u0394": 0.10000000000000003
1571
+ },
1572
+ {
1573
+ "task": "tev1_test/ag_news",
1574
+ "n": 120,
1575
+ "LightDec": 0.925,
1576
+ "proof_v2": 0.425,
1577
+ "\u0394": 0.5
1578
+ },
1579
+ {
1580
+ "task": "tev1_test/banking77",
1581
+ "n": 120,
1582
+ "LightDec": 0.7166666666666667,
1583
+ "proof_v2": 0.55,
1584
+ "\u0394": 0.16666666666666663
1585
+ },
1586
+ {
1587
+ "task": "tev1_test/boolq",
1588
+ "n": 120,
1589
+ "LightDec": 0.8416666666666667,
1590
+ "proof_v2": 0.5583333333333333,
1591
+ "\u0394": 0.2833333333333333
1592
+ },
1593
+ {
1594
+ "task": "tev1_test/mnli",
1595
+ "n": 120,
1596
+ "LightDec": 0.75,
1597
+ "proof_v2": 0.39166666666666666,
1598
+ "\u0394": 0.35833333333333334
1599
+ },
1600
+ {
1601
+ "task": "tev1_test/policy",
1602
+ "n": 120,
1603
+ "LightDec": 0.5,
1604
+ "proof_v2": 0.4083333333333333,
1605
+ "\u0394": 0.09166666666666667
1606
+ },
1607
+ {
1608
+ "task": "tev1_test/routing",
1609
+ "n": 120,
1610
+ "LightDec": 0.44166666666666665,
1611
+ "proof_v2": 0.2916666666666667,
1612
+ "\u0394": 0.14999999999999997
1613
+ },
1614
+ {
1615
+ "task": "tev1_test/sst5",
1616
+ "n": 120,
1617
+ "LightDec": 0.39166666666666666,
1618
+ "proof_v2": 0.275,
1619
+ "\u0394": 0.11666666666666664
1620
+ },
1621
+ {
1622
+ "task": "triage/support_email",
1623
+ "n": 120,
1624
+ "LightDec": 0.9583333333333334,
1625
+ "proof_v2": 0.39166666666666666,
1626
+ "\u0394": 0.5666666666666667
1627
+ },
1628
+ {
1629
+ "task": "typed_decisions/agent_trace_observability",
1630
+ "n": 120,
1631
+ "LightDec": 0.8083333333333333,
1632
+ "proof_v2": 0.3333333333333333,
1633
+ "\u0394": 0.47500000000000003
1634
+ },
1635
+ {
1636
+ "task": "typed_decisions/customer_service",
1637
+ "n": 120,
1638
+ "LightDec": 0.7416666666666667,
1639
+ "proof_v2": 0.2833333333333333,
1640
+ "\u0394": 0.45833333333333337
1641
+ },
1642
+ {
1643
+ "task": "typed_decisions/invoice_processing",
1644
+ "n": 120,
1645
+ "LightDec": 0.8166666666666667,
1646
+ "proof_v2": 0.475,
1647
+ "\u0394": 0.3416666666666667
1648
+ },
1649
+ {
1650
+ "task": "typed_decisions/security_incidents",
1651
+ "n": 120,
1652
+ "LightDec": 0.7583333333333333,
1653
+ "proof_v2": 0.5,
1654
+ "\u0394": 0.2583333333333333
1655
+ },
1656
+ {
1657
+ "task": "winogrande/blank",
1658
+ "n": 120,
1659
+ "LightDec": 0.6916666666666667,
1660
+ "proof_v2": 0.5583333333333333,
1661
+ "\u0394": 0.1333333333333333
1662
+ },
1663
+ {
1664
+ "task": "yelp/score",
1665
+ "n": 120,
1666
+ "LightDec": 0.625,
1667
+ "proof_v2": 0.325,
1668
+ "\u0394": 0.3
1669
+ }
1670
+ ],
1671
+ "head_to_head_error": null,
1672
+ "v3_benchmarks": {
1673
+ "tev1": {
1674
+ "per_task": [
1675
+ {
1676
+ "task": "tev1_test/ag_news",
1677
+ "n": 150,
1678
+ "acc": 0.92,
1679
+ "overlaps_LightDec_training_source": true
1680
+ },
1681
+ {
1682
+ "task": "tev1_test/banking77",
1683
+ "n": 200,
1684
+ "acc": 0.72,
1685
+ "overlaps_LightDec_training_source": false
1686
+ },
1687
+ {
1688
+ "task": "tev1_test/boolq",
1689
+ "n": 200,
1690
+ "acc": 0.835,
1691
+ "overlaps_LightDec_training_source": true
1692
+ },
1693
+ {
1694
+ "task": "tev1_test/mnli",
1695
+ "n": 300,
1696
+ "acc": 0.7233333333333334,
1697
+ "overlaps_LightDec_training_source": true
1698
+ },
1699
+ {
1700
+ "task": "tev1_test/policy",
1701
+ "n": 1200,
1702
+ "acc": 0.5291666666666667,
1703
+ "overlaps_LightDec_training_source": false
1704
+ },
1705
+ {
1706
+ "task": "tev1_test/routing",
1707
+ "n": 600,
1708
+ "acc": 0.4583333333333333,
1709
+ "overlaps_LightDec_training_source": false
1710
+ },
1711
+ {
1712
+ "task": "tev1_test/sst5",
1713
+ "n": 150,
1714
+ "acc": 0.3933333333333333,
1715
+ "overlaps_LightDec_training_source": false
1716
+ }
1717
+ ],
1718
+ "all": 0.5839285714285715,
1719
+ "no_overlap": 0.5176744186046511,
1720
+ "n": 2800,
1721
+ "published_tev1_4b": {
1722
+ "main decisions (1,000)": 0.88,
1723
+ "policy transfer (300)": 1.0
1724
+ }
1725
+ },
1726
+ "triage": {
1727
+ "emails": 101,
1728
+ "per_question": {
1729
+ "topic": 0.9900990099009901,
1730
+ "refund": 1.0,
1731
+ "breakage": 0.9900990099009901,
1732
+ "anger": 0.9207920792079208,
1733
+ "judgment": 0.8217821782178217
1734
+ },
1735
+ "pile_accuracy": 0.9405940594059405
1736
+ },
1737
+ "long_context": {
1738
+ "max_len": 2048,
1739
+ "per_task": [
1740
+ {
1741
+ "task": "long/contract_clause",
1742
+ "n": 150,
1743
+ "acc": 1.0,
1744
+ "chance": 0.19999999999999996
1745
+ },
1746
+ {
1747
+ "task": "long/email_thread",
1748
+ "n": 150,
1749
+ "acc": 1.0,
1750
+ "chance": 0.25
1751
+ },
1752
+ {
1753
+ "task": "long/service_log",
1754
+ "n": 150,
1755
+ "acc": 1.0,
1756
+ "chance": 0.19999999999999996
1757
+ }
1758
+ ]
1759
+ }
1760
+ },
1761
+ "architecture": "FalconDec",
1762
+ "latency": {
1763
+ "gpu": "NVIDIA RTX PRO 6000 Blackwell Server Edition",
1764
+ "calls": [
1765
+ {
1766
+ "device": "cuda",
1767
+ "weights": "fp16",
1768
+ "questions": 1,
1769
+ "p50_ms": 10.05,
1770
+ "p95_ms": 10.43
1771
+ },
1772
+ {
1773
+ "device": "cuda",
1774
+ "weights": "fp16",
1775
+ "questions": 5,
1776
+ "p50_ms": 11.32,
1777
+ "p95_ms": 11.41
1778
+ },
1779
+ {
1780
+ "device": "cuda",
1781
+ "weights": "fp16",
1782
+ "questions": 10,
1783
+ "p50_ms": 12.57,
1784
+ "p95_ms": 12.69
1785
+ },
1786
+ {
1787
+ "device": "cpu fp32",
1788
+ "weights": "int8 file",
1789
+ "questions": 1,
1790
+ "p50_ms": 48.4,
1791
+ "p95_ms": 48.8
1792
+ }
1793
+ ],
1794
+ "throughput_per_s": 2588.8502560238694
1795
+ },
1796
+ "env": {
1797
+ "torch": "2.9.0+cu130",
1798
+ "transformers": "5.17.0",
1799
+ "python": "3.12.12",
1800
+ "gpu": "NVIDIA RTX PRO 6000 Blackwell Server Edition"
1801
+ }
1802
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c675a38babe2bb34a81efc13936a134b7596cc5b12da859bf28966b537659e0b
3
+ size 319325890
tokenizer/tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer/tokenizer_config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": true,
4
+ "cls_token": "[CLS]",
5
+ "is_local": false,
6
+ "local_files_only": false,
7
+ "mask_token": "[MASK]",
8
+ "model_input_names": [
9
+ "input_ids",
10
+ "attention_mask"
11
+ ],
12
+ "model_max_length": 8192,
13
+ "pad_token": "[PAD]",
14
+ "sep_token": "[SEP]",
15
+ "tokenizer_class": "TokenizersBackend",
16
+ "unk_token": "[UNK]"
17
+ }