OpenDecider

OpenDecider-nano ONNX

opendecider-nano, the ~400M calibrated decision model, as ONNX for the browser (WebGPU or WebAssembly), Node, Bun and Deno. Give it a state (text or JSON) and typed questions (choice, score, noul); it returns a probability for every option, on the user's device: nothing is sent anywhere. These are the files @opendecider/web loads.

Try it in your browser · Guide · GitHub

npm install @opendecider/web
import { loadNano, choice, noul, Guard } from "@opendecider/web";

const model = await loadNano(); // these files, at the revision the package pins, SHA-256 checked, then cached
const r = await model.systemOne("I was charged twice for order 1182. Please fix this today.", {
  team: choice("Which team should handle this?", { billing: "charges, refunds", tech: "bugs, outages" }),
  urgent: noul("Does this need a reply today?"),
});
r.answers.team; // { type: "choice", choice: "billing", probabilities: { billing: 0.92, tech: 0.08 }, ... }

const guard = new Guard({ model }); // the prompt guard, on the device

Files

file build size SHA-256
onnx/model_q8f16.onnx 8-bit weights, float16 elsewhere: WebGPU's default 469,218,222 bytes (450 MiB) 80436dbed1a284548b5e6c3975bb26df81fb493fc4e64629e5a6ef490bf0e0ae
onnx/model_q8.onnx 8-bit weights, float32 elsewhere: WebAssembly's default 594,006,972 bytes (569 MiB) 3920323e41bb107f28566be7b45e6cb55a2bdc1951b27a6e62625857e22542c4
tokenizer.web.json opendecider-nano's tokenizer for tokenizers.js 1,570,653 bytes fd4146f0082d98fd1fc03bf41349f4d1eed478987ce9a64addc1654dfcad14ee

One graph each: input_ids and attention_mask (int64, [batch, tokens]) in, logits (float32, [batch, tokens]) out, one score per token; the answer is the softmax over the scores at the [MASK] markers, one before each option (opendecider.json has the details). The weights are 8-bit (ONNX Runtime's MatMulNBits, block 32): weight-only, so the browser's WebAssembly gives the same logits as native ONNX Runtime to 1e-6.

Same answers as the PyTorch model

Scored with native ONNX Runtime against opendecider-nano's committed answers, on every benchmark the model card reports:

build typed-decisions (2,000) general (200) Laya's battery (10 × 400) answers as PyTorch
opendecider-nano (PyTorch) 0.796 0.680 0.656 –
q8f16 0.798 0.680 0.656 99.5–99.8%
q8 0.795 0.680 0.656 99.6–99.8%

As a guard against jailbreaks and prompt injection (2,438 prompts, each build at its own train-split threshold, which @opendecider/web's Guard uses):

build threshold accuracy (95% CI) attacks caught benign flagged
opendecider-nano (PyTorch) 0.387 0.831 (0.816–0.845) 0.914 0.214
q8 0.391 0.828 (0.813–0.843) 0.914 0.218
q8f16 0.390 0.828 (0.813–0.842) 0.914 0.218

4-bit builds were not released: they changed 2.5–4.3% of the answers. Every number: Benchmarks.

Will it run on my device?

WebGPU, q8f16 WebAssembly, q8
download (once, then the browser's cache) 450 MiB 569 MiB
memory, with a 2,048-token question about 2.2 GiB about 2.8 GiB
one question, Chrome on an Apple M4 Max 47 ms 125 ms (8 threads), 830 ms (1 thread)

A browser with WebGPU uses it; otherwise WebAssembly. More WebAssembly threads need a cross-origin-isolated page (see the guide). Phones with little memory may not load it.

Without JavaScript

The graph runs in any ONNX Runtime. In Python, with the same prompt as the opendecider package (0.7.0 or later):

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
from opendecider.prompt import nano_ids

tok = AutoTokenizer.from_pretrained("manjunathshiva/opendecider-nano")
sess = ort.InferenceSession(hf_hub_download("manjunathshiva/opendecider-nano-ONNX", "onnx/model_q8.onnx"))
options = {"billing": "charges, refunds", "tech": "bugs, outages"}
ids, truncated = nano_ids(tok, "I was charged twice for order 1182.", "Which team should handle this?", options, 2048)
x = np.array([ids])
logits = sess.run(["logits"], {"input_ids": x, "attention_mask": np.ones_like(x)})[0][0]
z = logits[x[0] == tok.mask_token_id].astype(np.float64)
p = np.exp(z - z.max()); p /= p.sum()
dict(zip(options, p.round(4)))   # {'billing': 0.9476, 'tech': 0.0524}

Check the files yourself

The files rebuild byte for byte from the repository's packaging/onnx/ (torch 2.14.1, transformers 5.18.0, onnx 1.23.1, onnxscript 0.7.2, onnxruntime 1.30.0), from opendecider-nano at revision beeeb640f3333aec4ef78ed3b7bd0605f973a590:

python packaging/onnx/export_nano.py --out build/ --revision beeeb640f3333aec4ef78ed3b7bd0605f973a590
python packaging/onnx/quantize.py build/model.onnx build/onnx/
python packaging/onnx/make_web_tokenizer.py <opendecider-nano>/tokenizer.json build/tokenizer.web.json
shasum -a 256 build/onnx/*.onnx build/tokenizer.web.json

@opendecider/web pins this repository's revision and the size and SHA-256 of each file, and never uses a file that does not match.

Limits

  • English, like opendecider-nano: evaluated in English only.
  • 2,048 tokens per input; a longer state is shortened first (the answer is marked truncated).
  • Not for edge functions: Cloudflare Workers and similar runtimes have far less memory than the model needs.
  • opendecider-nano's own limits apply; see its model card.

License

Apache-2.0, as opendecider-nano. Built on Ettin-encoder-400m (MIT); see NOTICE. Manjunath Janardhan.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for manjunathshiva/opendecider-nano-ONNX

Quantized
(1)
this model

Collection including manjunathshiva/opendecider-nano-ONNX