divyanshudhruv commited on
Commit
2c91e82
·
verified ·
1 Parent(s): 1363687

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +235 -15
README.md CHANGED
@@ -1,15 +1,235 @@
1
- ---
2
- title: Oev Demo
3
- emoji: 📊
4
- colorFrom: green
5
- colorTo: indigo
6
- sdk: gradio
7
- sdk_version: 6.28.0
8
- python_version: '3.12'
9
- app_file: app.py
10
- pinned: false
11
- license: apache-2.0
12
- short_description: Typed questions in, Calibrated probabilities out
13
- ---
14
-
15
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <p align="center" style="margin: 24px 0;">
2
+ <picture>
3
+ <source media="(prefers-color-scheme: dark)" srcset="assets/logo_transp.png" />
4
+ <img src="assets/logo_transp.png" alt="OEV" width="150" />
5
+ </picture>
6
+ </p>
7
+
8
+ <div align="center">
9
+
10
+ OEV is a small (`184M params`) neural decision engine. Instead of generating text, it scores answer options directly. The state, the question, and every option are packed into one sequence. One forward pass returns a calibrated probability distribution.
11
+
12
+ The single `184M` model scores `0.7705` on typed-decisions, slightly above laya's published `0.766` from a `421M` checkpoint. An ensemble of four `184M` checkpoints reaches `0.7760`, the `highest` reported result, and a single Banking77 soup checkpoint reaches `0.8584` (best ECE `0.0595` from the 3-checkpoint ensemble). Jev leads only on Banking77 (0.870).
13
+
14
+ > [!WARNING]
15
+ > Chart latency comparisons use different hardware and include published ranges; the current `runs/` manifest and raw timing samples are also unavailable here. The sharpening panel in the comparison figure shows historical gamma `2.5` values from an evaluation sweep; they are exploratory and not release claims. Reproduce claims from recorded run logs before treating them as release evidence.
16
+
17
+ [![Hugging Face Model](https://img.shields.io/badge/%F0%9F%A4%97%20Model-divyanshudhruv%2Foev--typed-blue)](https://huggingface.co/divyanshudhruv/oev-typed)
18
+ [![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)
19
+ [![Tests](https://img.shields.io/badge/tests-passing-brightgreen)](https://github.com/divyanshudhruv/oev/actions/workflows/tests.yml)
20
+ [![Python](https://img.shields.io/badge/python-3.10%2B-blue)](https://www.python.org/downloads/)
21
+ [![PyTorch](https://img.shields.io/badge/PyTorch-2.x-ee4c2c)](https://pytorch.org/get-started/locally/)
22
+ [![HF Space](https://img.shields.io/badge/%F0%9F%A4%97%20Space-oev--demo-yellow)](https://huggingface.co/spaces/divyanshudhruv/oev-demo)
23
+ [![PyPI](https://img.shields.io/pypi/v/oev)](https://pypi.org/project/oev/)
24
+
25
+ </div>
26
+
27
+ <p align="center" style="margin: 24px 0;">
28
+ <picture>
29
+ <source media="(prefers-color-scheme: dark)" srcset="assets/benchmarks_dark.png" />
30
+ <img src="assets/benchmarks.png" alt="OEV vs Jev and laya on shared public benchmarks" width="92%" />
31
+ </picture>
32
+ </p>
33
+
34
+ <p align="center" style="margin: 24px 0;">
35
+ <picture>
36
+ <source media="(prefers-color-scheme: dark)" srcset="assets/oev_vs_jev_full_dark.png" />
37
+ <img src="assets/oev_vs_jev_full.png" alt="OEV versus TypeSafe Jev: accuracy on shared public datasets, every application workflow, speed, calibration, size, and soft-accuracy sharpening" width="95%" />
38
+ </picture>
39
+ </p>
40
+
41
+ <p align="center" style="margin: 24px 0;">
42
+ <picture>
43
+ <source media="(prefers-color-scheme: dark)" srcset="assets/transfer_speed_dark.png" />
44
+ <img src="assets/transfer_speed.png" alt="Left: zero-shot and out-of-domain transfer, OEV distilled student beats laya zero-shot on emotion with the NLI floor documented; right: latency, OEV 22.2ms on T4 vs laya, Kev-4B and Jev" width="100%" />
45
+ </picture>
46
+ </p>
47
+
48
+ ## At a glance
49
+
50
+ | claim | result |
51
+ | --------------------- | ------------------------------------------------------------------------------------------------------- |
52
+ | best accuracy | **0.7705** single model, **0.7760** ensemble - typed-decisions (laya 0.766 from 421M) |
53
+ | high-cardinality | **0.8584** Banking77, one soup checkpoint (`ECE 0.0595` best ensemble; laya 0.425) |
54
+ | speed | **22.2 ms** single question (laya 32.8-39.5 ms published range) |
55
+ | size | **184M** params, 0.44x laya |
56
+ | weights & checkpoints | Apache 2.0 - [huggingface.co/divyanshudhruv/oev-typed](https://huggingface.co/divyanshudhruv/oev-typed) |
57
+
58
+ - `22.2 ms` per question on a `T4` (GPU); `447 ms` p50 on CPU (8 threads, 184M soup checkpoint)
59
+ - `0.8584` on 77-label `Banking77` from a single soup checkpoint (the 3-checkpoint ensemble still holds best ECE `0.0595`): each option is embedded as its own anchor with full tokens, so accuracy scales with label count (gap to Jev 1.16 pts)
60
+ - `184M` params, `Apache 2.0` weights
61
+ - Kev (0.8B / 4B) publishes no in-domain numbers on these datasets, so it is not in the tables; see [BENCHMARKS.md](BENCHMARKS.md) for the like-for-like comparison plan
62
+
63
+
64
+ ## Architecture
65
+
66
+ ```mermaid
67
+ flowchart LR
68
+ S["state\n(text / JSON)"] --> P["packer:\nstate + questions + anchors\none sequence"]
69
+ P --> E["encoder\nDeBERTa-v3-base (184M)\nor char transformer"]
70
+ E --> H["one linear head\nscores every ANCHOR"]
71
+ H --> D["softmax per question\n= calibrated distribution"]
72
+ D --> O["choice / noul / score"]
73
+ ```
74
+
75
+ - **One anchor mechanism** covers all three primitives - options, yes/no pairs, and score levels are each embedded as anchors in one packed sequence
76
+ - **No text generation** - nothing to parse, nothing to hallucinate
77
+
78
+ > **New question types need no new heads**
79
+
80
+ ## Benchmarks: OEV vs the published field
81
+
82
+ Fine-tuned on each benchmark's train split, following the same protocol as Laya's published runs. Selected results and evaluation notes are in [BENCHMARKS.md](BENCHMARKS.md).
83
+
84
+ > [!WARNING]
85
+ > Banking77 uses 77 OEV labels, while the published Jev figure is from a 72-label configuration. The `0.8584` and `0.870` values are not a controlled head-to-head comparison.
86
+
87
+ | benchmark | OEV | laya | Jev | note |
88
+ | --------------- | ---------: | ----: | ----: | ----------------------------------------------------------------- |
89
+ | typed-decisions | **0.7760** | 0.766 | 0.727 | highest reported (ensemble); single model 0.7705 |
90
+ | Banking77 | **0.8584** | 0.425 | 0.870 | 2x laya; gap to Jev 1.16 pts |
91
+ | AG News | **0.9489** | 0.950 | 0.910 | label-noise ceiling (~0.95) |
92
+ | DAIR Emotion | **0.9300** | 0.595 | 0.480 | zero-shot: OEV student `0.6505` beats laya's `0.595` head-to-head |
93
+
94
+ <p align="center" style="margin: 24px 0;">
95
+ <picture>
96
+ <source media="(prefers-color-scheme: dark)" srcset="assets/headline_scorecard_dark.png" />
97
+ <img src="assets/headline_scorecard.png" alt="OEV headline results: typed-decisions accuracy, Banking77 accuracy, and hardware-separated latency" width="100%" />
98
+ </picture>
99
+ </p>
100
+
101
+ <p align="center" style="margin: 24px 0;">
102
+ <picture>
103
+ <source media="(prefers-color-scheme: dark)" srcset="assets/zeroshot_dark.png" />
104
+ <img src="assets/zeroshot.png" alt="Zero-shot and out-of-domain transfer as dot pairs: OEV 0.650 versus laya 0.595 on emotion, with the WANLI and ANLI floors marked" width="49%" />
105
+ </picture>
106
+ <picture>
107
+ <source media="(prefers-color-scheme: dark)" srcset="assets/latency_profile_dark.png" />
108
+ <img src="assets/latency_profile.png" alt="OEV latency on Tesla T4 and CPU, hardware separated; batch value is per-question throughput" width="49%" />
109
+ </picture>
110
+ </p>
111
+
112
+ <p align="center" style="margin: 24px 0;">
113
+
114
+ <picture>
115
+ <source media="(prefers-color-scheme: dark)" srcset="assets/workflows_dark.png" />
116
+ <img src="assets/workflows.png" alt="Typed-decisions accuracy per workflow: OEV wins invoice processing, customer service and agent-trace observability; laya wins security incidents" width="49%" />
117
+ </picture><picture>
118
+ <source media="(prefers-color-scheme: dark)" srcset="assets/decision_primitives_dark.png" />
119
+ <img src="assets/decision_primitives.png" alt="Illustrative normalized distributions for OEV choice, noul, and score decision primitives" width="50%" />
120
+ </picture>
121
+ </p>
122
+
123
+ ## Quickstart
124
+
125
+ ```bash
126
+ pip install oev # from PyPI
127
+ # or from source:
128
+ pip install -e .
129
+ ```
130
+
131
+ Optional extras:
132
+
133
+ - `pip install -e ".[dev]"` - pytest
134
+ - `pip install -e ".[data]"` - dataset converters
135
+ - `pip install -e ".[backbone]"` - DeBERTa fine-tuning
136
+ - `pip install -e ".[app]"` - Gradio Space dependencies
137
+
138
+ ```python
139
+ from oev.infer import OEV
140
+
141
+ agent = OEV("checkpoints_td5/oev-tiny.pt", device="cpu")
142
+
143
+ result = agent.decide("We were charged twice for the same order.", {
144
+ "department": {"type": "choice", "options": ["billing", "technical", "sales", "other"],
145
+ "instructions": "Which department should handle this?"},
146
+ "refund_requested": {"type": "noul"},
147
+ "severity": {"type": "score", "levels": [1, 2, 3, 4, 5]},
148
+ })
149
+ ```
150
+
151
+ ```json
152
+ {
153
+ "department": {
154
+ "choice": "billing",
155
+ "probabilities": {
156
+ "billing": 0.94,
157
+ "technical": 0.04,
158
+ "sales": 0.01,
159
+ "other": 0.01
160
+ },
161
+ "confidence": 0.94
162
+ },
163
+ "refund_requested": 0.91,
164
+ "severity": {
165
+ "value": 3,
166
+ "probabilities": { "1": 0.02, "2": 0.08, "3": 0.61, "4": 0.22, "5": 0.07 }
167
+ }
168
+ }
169
+ ```
170
+
171
+ _Confidence gating_ - automate when confident, escalate when not:
172
+
173
+ ```python
174
+ from oev.presets import triage_questions, gate
175
+
176
+ for name, payload, confident in gate(result, threshold=0.85):
177
+ automate(name, payload) if confident else escalate_to_human(name)
178
+ ```
179
+
180
+ HTTP server (native + Jev-compatible `/v1/systemone` endpoint - TypeSafe clients work by changing baseUrl):
181
+
182
+ ```bash
183
+ pip install -e ".[serve]"
184
+ oev-serve --checkpoint checkpoints_td5/oev-tiny.pt --port 8000
185
+ curl -X POST localhost:8000/decide -H "Content-Type: application/json" \
186
+ -d '{"state": "My payment failed twice", "questions": {"urgency": {"type": "score", "levels": [1, 2, 3]}}}'
187
+ ```
188
+
189
+ Docker:
190
+
191
+ ```bash
192
+ docker compose up # checkpoint at ./checkpoints/oev-tiny.pt
193
+ ```
194
+
195
+ ## Training
196
+
197
+ ```bash
198
+ python -m oev.convert_typed
199
+
200
+ python -m oev.train --backbone microsoft/deberta-v3-base --epochs 4 --batch-size 8 --max-len 768 --data-dir data/typed --out checkpoints_td5
201
+
202
+ python -m oev.rlcd --checkpoint checkpoints_td5/oev-tiny.pt --data-dir data/typed --epochs 2 --batch-size 8 --out checkpoints_rlcd
203
+
204
+ python -m oev.ensemble --ckpts checkpoints_td5/oev-tiny.pt,checkpoints_rlcd/oev-tiny.pt,checkpoints_rlcd_soup/oev-tiny.pt,checkpoints_rlcd_seed1/oev-tiny.pt --data-dir data/typed
205
+
206
+ python -m pytest -q
207
+ ```
208
+
209
+ `train_colab.ipynb` runs the entire pipeline end to end. Evaluation rules and selected limitations are in [BENCHMARKS.md](BENCHMARKS.md).
210
+
211
+ ## Limitations
212
+
213
+ - every benchmark number is from a checkpoint fine-tuned on that benchmark's train split. Zero-shot emotion is a **win**: the shipped distilled student (`0.6505`, both models zero-shot) against laya's `0.595`, up from the round-1 starting point of `0.4265`
214
+ - ANLI remains near chance (`0.3380` for the round-1 student), while the WANLI specialist reaches `0.5645`; the NLI result is split-dependent, not uniformly at chance
215
+ - pure-mimicry distillation transfers breadth, not depth. The round-1 student hit emotion `0.6875` (+26 pts zero-shot) but lost typed skill (`0.5385`). Adding gold-CE loss (round 2) collapsed to uniform, a documented negative result. Round 2b (pure-KL, balanced domains) rescued it: typed `0.6480`, emotion `0.6505`, b77 `0.7964`, probes all PASS
216
+ - the headline ensemble result is an average of four checkpoints; the best single model is `0.7705`
217
+ - CPU inference is roughly `20x` slower than the T4 (`447 ms` p50, 8 threads) - all headline timings are GPU
218
+ - the b77 headline includes a 3-checkpoint probability ensemble and a single-file soup checkpoint at `0.8584`; `0.8403` is a historical warm-start re-tune, not the current best single artifact
219
+ - English only
220
+
221
+ ## Roadmap
222
+
223
+ - [ ] round 3 distillation: 6 teachers, 5 domains including NLI
224
+ - [ ] 4-member Banking77 ensemble: the live shot past `0.8584`
225
+ - [ ] round 4: one file near specialist numbers everywhere
226
+ - [ ] INT8 / ONNX CPU deployment (export + quantization scripts in `scripts/`, bench pending)
227
+ - [ ] multi-question shared-state encoding (one pass, many questions)
228
+ - [ ] robustness: reduce mild overconfidence on out-of-distribution inputs
229
+ - [ ] non-English checkpoints (the interface is language-agnostic; the weights are not yet)
230
+
231
+ Full list: [ROADMAP.md](ROADMAP.md).
232
+
233
+ ## Credits
234
+
235
+ The interface and benchmark protocol follow [Laya](https://github.com/NandhaKishorM/laya), [Kev](https://github.com/jaredpalmer/kev) and the System One model category introduced by TypeSafe's [Jev](https://typesafe.com). Their published numbers are quoted here for comparison and remain their measurements.