Spaces:
Running on Zero
Running on Zero
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,15 +1,235 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
<p align="center" style="margin: 24px 0;">
|
| 2 |
+
<picture>
|
| 3 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/logo_transp.png" />
|
| 4 |
+
<img src="assets/logo_transp.png" alt="OEV" width="150" />
|
| 5 |
+
</picture>
|
| 6 |
+
</p>
|
| 7 |
+
|
| 8 |
+
<div align="center">
|
| 9 |
+
|
| 10 |
+
OEV is a small (`184M params`) neural decision engine. Instead of generating text, it scores answer options directly. The state, the question, and every option are packed into one sequence. One forward pass returns a calibrated probability distribution.
|
| 11 |
+
|
| 12 |
+
The single `184M` model scores `0.7705` on typed-decisions, slightly above laya's published `0.766` from a `421M` checkpoint. An ensemble of four `184M` checkpoints reaches `0.7760`, the `highest` reported result, and a single Banking77 soup checkpoint reaches `0.8584` (best ECE `0.0595` from the 3-checkpoint ensemble). Jev leads only on Banking77 (0.870).
|
| 13 |
+
|
| 14 |
+
> [!WARNING]
|
| 15 |
+
> Chart latency comparisons use different hardware and include published ranges; the current `runs/` manifest and raw timing samples are also unavailable here. The sharpening panel in the comparison figure shows historical gamma `2.5` values from an evaluation sweep; they are exploratory and not release claims. Reproduce claims from recorded run logs before treating them as release evidence.
|
| 16 |
+
|
| 17 |
+
[](https://huggingface.co/divyanshudhruv/oev-typed)
|
| 18 |
+
[](LICENSE)
|
| 19 |
+
[](https://github.com/divyanshudhruv/oev/actions/workflows/tests.yml)
|
| 20 |
+
[](https://www.python.org/downloads/)
|
| 21 |
+
[](https://pytorch.org/get-started/locally/)
|
| 22 |
+
[](https://huggingface.co/spaces/divyanshudhruv/oev-demo)
|
| 23 |
+
[](https://pypi.org/project/oev/)
|
| 24 |
+
|
| 25 |
+
</div>
|
| 26 |
+
|
| 27 |
+
<p align="center" style="margin: 24px 0;">
|
| 28 |
+
<picture>
|
| 29 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/benchmarks_dark.png" />
|
| 30 |
+
<img src="assets/benchmarks.png" alt="OEV vs Jev and laya on shared public benchmarks" width="92%" />
|
| 31 |
+
</picture>
|
| 32 |
+
</p>
|
| 33 |
+
|
| 34 |
+
<p align="center" style="margin: 24px 0;">
|
| 35 |
+
<picture>
|
| 36 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/oev_vs_jev_full_dark.png" />
|
| 37 |
+
<img src="assets/oev_vs_jev_full.png" alt="OEV versus TypeSafe Jev: accuracy on shared public datasets, every application workflow, speed, calibration, size, and soft-accuracy sharpening" width="95%" />
|
| 38 |
+
</picture>
|
| 39 |
+
</p>
|
| 40 |
+
|
| 41 |
+
<p align="center" style="margin: 24px 0;">
|
| 42 |
+
<picture>
|
| 43 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/transfer_speed_dark.png" />
|
| 44 |
+
<img src="assets/transfer_speed.png" alt="Left: zero-shot and out-of-domain transfer, OEV distilled student beats laya zero-shot on emotion with the NLI floor documented; right: latency, OEV 22.2ms on T4 vs laya, Kev-4B and Jev" width="100%" />
|
| 45 |
+
</picture>
|
| 46 |
+
</p>
|
| 47 |
+
|
| 48 |
+
## At a glance
|
| 49 |
+
|
| 50 |
+
| claim | result |
|
| 51 |
+
| --------------------- | ------------------------------------------------------------------------------------------------------- |
|
| 52 |
+
| best accuracy | **0.7705** single model, **0.7760** ensemble - typed-decisions (laya 0.766 from 421M) |
|
| 53 |
+
| high-cardinality | **0.8584** Banking77, one soup checkpoint (`ECE 0.0595` best ensemble; laya 0.425) |
|
| 54 |
+
| speed | **22.2 ms** single question (laya 32.8-39.5 ms published range) |
|
| 55 |
+
| size | **184M** params, 0.44x laya |
|
| 56 |
+
| weights & checkpoints | Apache 2.0 - [huggingface.co/divyanshudhruv/oev-typed](https://huggingface.co/divyanshudhruv/oev-typed) |
|
| 57 |
+
|
| 58 |
+
- `22.2 ms` per question on a `T4` (GPU); `447 ms` p50 on CPU (8 threads, 184M soup checkpoint)
|
| 59 |
+
- `0.8584` on 77-label `Banking77` from a single soup checkpoint (the 3-checkpoint ensemble still holds best ECE `0.0595`): each option is embedded as its own anchor with full tokens, so accuracy scales with label count (gap to Jev 1.16 pts)
|
| 60 |
+
- `184M` params, `Apache 2.0` weights
|
| 61 |
+
- Kev (0.8B / 4B) publishes no in-domain numbers on these datasets, so it is not in the tables; see [BENCHMARKS.md](BENCHMARKS.md) for the like-for-like comparison plan
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
## Architecture
|
| 65 |
+
|
| 66 |
+
```mermaid
|
| 67 |
+
flowchart LR
|
| 68 |
+
S["state\n(text / JSON)"] --> P["packer:\nstate + questions + anchors\none sequence"]
|
| 69 |
+
P --> E["encoder\nDeBERTa-v3-base (184M)\nor char transformer"]
|
| 70 |
+
E --> H["one linear head\nscores every ANCHOR"]
|
| 71 |
+
H --> D["softmax per question\n= calibrated distribution"]
|
| 72 |
+
D --> O["choice / noul / score"]
|
| 73 |
+
```
|
| 74 |
+
|
| 75 |
+
- **One anchor mechanism** covers all three primitives - options, yes/no pairs, and score levels are each embedded as anchors in one packed sequence
|
| 76 |
+
- **No text generation** - nothing to parse, nothing to hallucinate
|
| 77 |
+
|
| 78 |
+
> **New question types need no new heads**
|
| 79 |
+
|
| 80 |
+
## Benchmarks: OEV vs the published field
|
| 81 |
+
|
| 82 |
+
Fine-tuned on each benchmark's train split, following the same protocol as Laya's published runs. Selected results and evaluation notes are in [BENCHMARKS.md](BENCHMARKS.md).
|
| 83 |
+
|
| 84 |
+
> [!WARNING]
|
| 85 |
+
> Banking77 uses 77 OEV labels, while the published Jev figure is from a 72-label configuration. The `0.8584` and `0.870` values are not a controlled head-to-head comparison.
|
| 86 |
+
|
| 87 |
+
| benchmark | OEV | laya | Jev | note |
|
| 88 |
+
| --------------- | ---------: | ----: | ----: | ----------------------------------------------------------------- |
|
| 89 |
+
| typed-decisions | **0.7760** | 0.766 | 0.727 | highest reported (ensemble); single model 0.7705 |
|
| 90 |
+
| Banking77 | **0.8584** | 0.425 | 0.870 | 2x laya; gap to Jev 1.16 pts |
|
| 91 |
+
| AG News | **0.9489** | 0.950 | 0.910 | label-noise ceiling (~0.95) |
|
| 92 |
+
| DAIR Emotion | **0.9300** | 0.595 | 0.480 | zero-shot: OEV student `0.6505` beats laya's `0.595` head-to-head |
|
| 93 |
+
|
| 94 |
+
<p align="center" style="margin: 24px 0;">
|
| 95 |
+
<picture>
|
| 96 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/headline_scorecard_dark.png" />
|
| 97 |
+
<img src="assets/headline_scorecard.png" alt="OEV headline results: typed-decisions accuracy, Banking77 accuracy, and hardware-separated latency" width="100%" />
|
| 98 |
+
</picture>
|
| 99 |
+
</p>
|
| 100 |
+
|
| 101 |
+
<p align="center" style="margin: 24px 0;">
|
| 102 |
+
<picture>
|
| 103 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/zeroshot_dark.png" />
|
| 104 |
+
<img src="assets/zeroshot.png" alt="Zero-shot and out-of-domain transfer as dot pairs: OEV 0.650 versus laya 0.595 on emotion, with the WANLI and ANLI floors marked" width="49%" />
|
| 105 |
+
</picture>
|
| 106 |
+
<picture>
|
| 107 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/latency_profile_dark.png" />
|
| 108 |
+
<img src="assets/latency_profile.png" alt="OEV latency on Tesla T4 and CPU, hardware separated; batch value is per-question throughput" width="49%" />
|
| 109 |
+
</picture>
|
| 110 |
+
</p>
|
| 111 |
+
|
| 112 |
+
<p align="center" style="margin: 24px 0;">
|
| 113 |
+
|
| 114 |
+
<picture>
|
| 115 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/workflows_dark.png" />
|
| 116 |
+
<img src="assets/workflows.png" alt="Typed-decisions accuracy per workflow: OEV wins invoice processing, customer service and agent-trace observability; laya wins security incidents" width="49%" />
|
| 117 |
+
</picture><picture>
|
| 118 |
+
<source media="(prefers-color-scheme: dark)" srcset="assets/decision_primitives_dark.png" />
|
| 119 |
+
<img src="assets/decision_primitives.png" alt="Illustrative normalized distributions for OEV choice, noul, and score decision primitives" width="50%" />
|
| 120 |
+
</picture>
|
| 121 |
+
</p>
|
| 122 |
+
|
| 123 |
+
## Quickstart
|
| 124 |
+
|
| 125 |
+
```bash
|
| 126 |
+
pip install oev # from PyPI
|
| 127 |
+
# or from source:
|
| 128 |
+
pip install -e .
|
| 129 |
+
```
|
| 130 |
+
|
| 131 |
+
Optional extras:
|
| 132 |
+
|
| 133 |
+
- `pip install -e ".[dev]"` - pytest
|
| 134 |
+
- `pip install -e ".[data]"` - dataset converters
|
| 135 |
+
- `pip install -e ".[backbone]"` - DeBERTa fine-tuning
|
| 136 |
+
- `pip install -e ".[app]"` - Gradio Space dependencies
|
| 137 |
+
|
| 138 |
+
```python
|
| 139 |
+
from oev.infer import OEV
|
| 140 |
+
|
| 141 |
+
agent = OEV("checkpoints_td5/oev-tiny.pt", device="cpu")
|
| 142 |
+
|
| 143 |
+
result = agent.decide("We were charged twice for the same order.", {
|
| 144 |
+
"department": {"type": "choice", "options": ["billing", "technical", "sales", "other"],
|
| 145 |
+
"instructions": "Which department should handle this?"},
|
| 146 |
+
"refund_requested": {"type": "noul"},
|
| 147 |
+
"severity": {"type": "score", "levels": [1, 2, 3, 4, 5]},
|
| 148 |
+
})
|
| 149 |
+
```
|
| 150 |
+
|
| 151 |
+
```json
|
| 152 |
+
{
|
| 153 |
+
"department": {
|
| 154 |
+
"choice": "billing",
|
| 155 |
+
"probabilities": {
|
| 156 |
+
"billing": 0.94,
|
| 157 |
+
"technical": 0.04,
|
| 158 |
+
"sales": 0.01,
|
| 159 |
+
"other": 0.01
|
| 160 |
+
},
|
| 161 |
+
"confidence": 0.94
|
| 162 |
+
},
|
| 163 |
+
"refund_requested": 0.91,
|
| 164 |
+
"severity": {
|
| 165 |
+
"value": 3,
|
| 166 |
+
"probabilities": { "1": 0.02, "2": 0.08, "3": 0.61, "4": 0.22, "5": 0.07 }
|
| 167 |
+
}
|
| 168 |
+
}
|
| 169 |
+
```
|
| 170 |
+
|
| 171 |
+
_Confidence gating_ - automate when confident, escalate when not:
|
| 172 |
+
|
| 173 |
+
```python
|
| 174 |
+
from oev.presets import triage_questions, gate
|
| 175 |
+
|
| 176 |
+
for name, payload, confident in gate(result, threshold=0.85):
|
| 177 |
+
automate(name, payload) if confident else escalate_to_human(name)
|
| 178 |
+
```
|
| 179 |
+
|
| 180 |
+
HTTP server (native + Jev-compatible `/v1/systemone` endpoint - TypeSafe clients work by changing baseUrl):
|
| 181 |
+
|
| 182 |
+
```bash
|
| 183 |
+
pip install -e ".[serve]"
|
| 184 |
+
oev-serve --checkpoint checkpoints_td5/oev-tiny.pt --port 8000
|
| 185 |
+
curl -X POST localhost:8000/decide -H "Content-Type: application/json" \
|
| 186 |
+
-d '{"state": "My payment failed twice", "questions": {"urgency": {"type": "score", "levels": [1, 2, 3]}}}'
|
| 187 |
+
```
|
| 188 |
+
|
| 189 |
+
Docker:
|
| 190 |
+
|
| 191 |
+
```bash
|
| 192 |
+
docker compose up # checkpoint at ./checkpoints/oev-tiny.pt
|
| 193 |
+
```
|
| 194 |
+
|
| 195 |
+
## Training
|
| 196 |
+
|
| 197 |
+
```bash
|
| 198 |
+
python -m oev.convert_typed
|
| 199 |
+
|
| 200 |
+
python -m oev.train --backbone microsoft/deberta-v3-base --epochs 4 --batch-size 8 --max-len 768 --data-dir data/typed --out checkpoints_td5
|
| 201 |
+
|
| 202 |
+
python -m oev.rlcd --checkpoint checkpoints_td5/oev-tiny.pt --data-dir data/typed --epochs 2 --batch-size 8 --out checkpoints_rlcd
|
| 203 |
+
|
| 204 |
+
python -m oev.ensemble --ckpts checkpoints_td5/oev-tiny.pt,checkpoints_rlcd/oev-tiny.pt,checkpoints_rlcd_soup/oev-tiny.pt,checkpoints_rlcd_seed1/oev-tiny.pt --data-dir data/typed
|
| 205 |
+
|
| 206 |
+
python -m pytest -q
|
| 207 |
+
```
|
| 208 |
+
|
| 209 |
+
`train_colab.ipynb` runs the entire pipeline end to end. Evaluation rules and selected limitations are in [BENCHMARKS.md](BENCHMARKS.md).
|
| 210 |
+
|
| 211 |
+
## Limitations
|
| 212 |
+
|
| 213 |
+
- every benchmark number is from a checkpoint fine-tuned on that benchmark's train split. Zero-shot emotion is a **win**: the shipped distilled student (`0.6505`, both models zero-shot) against laya's `0.595`, up from the round-1 starting point of `0.4265`
|
| 214 |
+
- ANLI remains near chance (`0.3380` for the round-1 student), while the WANLI specialist reaches `0.5645`; the NLI result is split-dependent, not uniformly at chance
|
| 215 |
+
- pure-mimicry distillation transfers breadth, not depth. The round-1 student hit emotion `0.6875` (+26 pts zero-shot) but lost typed skill (`0.5385`). Adding gold-CE loss (round 2) collapsed to uniform, a documented negative result. Round 2b (pure-KL, balanced domains) rescued it: typed `0.6480`, emotion `0.6505`, b77 `0.7964`, probes all PASS
|
| 216 |
+
- the headline ensemble result is an average of four checkpoints; the best single model is `0.7705`
|
| 217 |
+
- CPU inference is roughly `20x` slower than the T4 (`447 ms` p50, 8 threads) - all headline timings are GPU
|
| 218 |
+
- the b77 headline includes a 3-checkpoint probability ensemble and a single-file soup checkpoint at `0.8584`; `0.8403` is a historical warm-start re-tune, not the current best single artifact
|
| 219 |
+
- English only
|
| 220 |
+
|
| 221 |
+
## Roadmap
|
| 222 |
+
|
| 223 |
+
- [ ] round 3 distillation: 6 teachers, 5 domains including NLI
|
| 224 |
+
- [ ] 4-member Banking77 ensemble: the live shot past `0.8584`
|
| 225 |
+
- [ ] round 4: one file near specialist numbers everywhere
|
| 226 |
+
- [ ] INT8 / ONNX CPU deployment (export + quantization scripts in `scripts/`, bench pending)
|
| 227 |
+
- [ ] multi-question shared-state encoding (one pass, many questions)
|
| 228 |
+
- [ ] robustness: reduce mild overconfidence on out-of-distribution inputs
|
| 229 |
+
- [ ] non-English checkpoints (the interface is language-agnostic; the weights are not yet)
|
| 230 |
+
|
| 231 |
+
Full list: [ROADMAP.md](ROADMAP.md).
|
| 232 |
+
|
| 233 |
+
## Credits
|
| 234 |
+
|
| 235 |
+
The interface and benchmark protocol follow [Laya](https://github.com/NandhaKishorM/laya), [Kev](https://github.com/jaredpalmer/kev) and the System One model category introduced by TypeSafe's [Jev](https://typesafe.com). Their published numbers are quoted here for comparison and remain their measurements.
|