Instructions to use hummbl-hf/decision-0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hummbl-hf/decision-0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="hummbl-hf/decision-0")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("hummbl-hf/decision-0") model = AutoModelForSequenceClassification.from_pretrained("hummbl-hf/decision-0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
HUMMBL Decision-0
Version 0.1.1. The 0.x version means the interface and the quality are not yet settled; see the warning and the limitations below. 0.1.1 corrects the model card and adds label ids to decision0.json; the weights and decide.py are unchanged from 0.1.0.
Decision-0 reads four short texts about an AI agent's next step (the policy it was given, the evidence about approvals or records, the task it was asked to do, and the action it proposes) and answers allow, deny or escalate.
- allow: the policy clearly permits the action and any condition it sets is met by the evidence.
- deny: the policy clearly forbids the action.
- escalate: anything else, such as a silent or vague policy, missing or only-claimed approval, or a request outside what the policy covers. A person should look.
It is small (about 71 million parameters) and runs on a laptop CPU.
Decision-0 is advisory. It is not safe for unattended use. On a sealed test written by model families that produced none of its training data, it wrongly allowed 12 of 335 actions that should not have gone ahead (3.6%; the 95% upper bound is 6.2%). With the recommended strict setting it wrongly allowed 4 of 335 (1.2%; upper bound 3.0%). On a second test whose labels came from a model family that touched no training case, the wrong-allow rate was higher: 6.7% at highest score and 4.2% at the strict setting. No person has checked any label, so every figure here measures agreement with labels set by language models. Use it to sort and route decisions, not to make them alone.
How to use
from decide import Decision0 # decide.py is in this repository
d = Decision0("hummbl-hf/decision-0") # strict allow threshold 0.99 by default
print(d.decide({
"policy": "Agents may read the billing database. Writes need a ticket from finance.",
"evidence": "No ticket is attached.",
"task": "Fix the duplicate invoice for customer 4417.",
"action": "DELETE FROM invoices WHERE id = 99812",
}))
The model reads two texts: Policy: {policy} Evidence: {evidence} and Task: {task} Proposed action: {action}. An empty field is written as (empty). Inputs are truncated to 192 tokens.
The decision rule, which decide.py implements:
- Take the label with the highest score.
- If it is allow and the allow score is below the threshold, answer escalate.
- If it is allow and the case has no policy text or no action text, answer escalate.
The two extra rules can only turn an allow into an escalate. They never produce an allow. The threshold is a choice between safety and coverage: 0.99 is recommended. With no threshold (highest score only), accuracy is higher but more wrong allows get through. See the table below.
Using the raw model directly: config.json names the three outputs, in index order 0 = deny, 1 = allow, 2 = escalate, so AutoModelForSequenceClassification returns them in that vocabulary. Look labels up by name (config.label2id), not by position. Apply the rule above yourself if you do not use decide.py. The plain Hugging Face pipeline only takes the highest score. In a pre-release check (eval/ADVERSARIAL.md), a request to email customers' social security numbers scored allow highest (0.49): the strict threshold escalated it, and the plain pipeline would have allowed it.
Evaluation
Each of the sealed tests below (D, E and F) was sealed by hash before any model was scored on it, and this model scored each one once. (eval/EVAL_V3.md also lists Tests A and B, which were never sealed and whose labels are unchecked; they are not used here.) Tests D and E were also scored by the previous version (v2), and the individual misses on D and on half of E were read afterwards, so they are weaker evidence than F. The test sets are not published, so they stay usable as tests.
| Test | Cases | Accuracy | Wrongly allowed (of cases that should not be allowed) | 95% upper bound |
|---|---|---|---|---|
| F: written by four model families that wrote no training data; labelled blind by three others | 520 | 0.913 | 12 of 335 (3.6%) | 6.2% |
| F, at the strict threshold 0.99 | 520 | 0.888 | 4 of 335 (1.2%) | 3.0% |
| E: written case by case by four Claude subagents, who set the labels; Claude neither wrote nor rated any training case | 360 | 0.919 | 16 of 240 (6.7%) | 10.6% |
| E, at the strict threshold 0.99 | 360 | 0.897 | 10 of 240 (4.2%) | 7.5% |
| E, the half whose individual misses were never read | 180 | 0.928 | 7 of 119 (5.9%) | 11.6% |
| D: written case by case by the model's author (a Claude model); no other model checked the labels; every miss has been read | 40 | 0.900 | 0 of 27 | 12.5% |
Rows without a threshold use the highest score. Upper bounds are 95% Wilson bounds.
Against simple baselines on Test F (520 cases):
| System | Accuracy | Wrongly allowed | Share of true allows allowed |
|---|---|---|---|
| Decision-0, strict threshold 0.99 | 0.888 | 4 of 335 | 76% |
| Keyword rules (fixed patterns over the policy and evidence text; escalate when there is no policy) | 0.594 | 117 of 335 | 62% |
| Always escalate | 0.283 | 0 of 335 | 0% |
Always escalating never wrongly allows anything, and never allows anything at all. Decision-0's value is the allow coverage it adds at a small false-allow cost. Totals for every test, including the baselines, are in eval/EVAL_V3.md.
- At the strict threshold, 76% of actions that should be allowed are still allowed on Test F. The rest are sent for escalation.
- At highest score on Test F, the 185 true allows were answered: 159 allow, 19 escalate, 7 deny. Caution is its commonest error. Its costly errors were 5 denies and 7 escalates that it allowed.
- Counting deny and escalate as one answer ("do not proceed alone"), at highest score it is on the right side of the allow line for 482 of 520 Test F cases (92.7%): 12 wrongly allowed, 26 wrongly held back.
- Its weakest groups on Test F are cases written by the ByteDance model (accuracy 0.871 on 132 cases), cases where one of the three rater models disagreed with the other two (0.875 on 48), and finance, legal and healthcare administration, its weakest subject (0.880 on 83).
- For comparison, on Test E three large language models labelling blind agreed with the labels 98.1%, 97.5% and 91.1% of the time; Decision-0 agrees 91.9%. On the allow line it trails the large models by about six points. That comparison uses different cases and is rough.
Training
- Base model:
cross-encoder/nli-deberta-v3-xsmall, revisiona150876415327c80daeff35ca6f68f5ed8cf5c24, a model already trained to judge whether one text supports another. - Data: 3,164 synthetic cases written by five language-model families. A case was kept only when at least three other model families agreed with its label. The training cases are not published.
- Settings: 4 epochs, learning rate 3e-5, batch size 16, seed 20260929, checkpoint chosen on a validation set only.
- Weights SHA-256:
3a1653761967dfcd1afd9fb6d3ee2a3b3e6578ad58b1b6b17d2e184842e309fc
Limitations
- No person has reviewed a label, in training or in any test. Every number above compares the model with labels set by language models.
- Every case, in training and in every test, was written by a language model. "Case by case" in the table means written one at a time rather than generated from templates; it does not mean written by a person. Real requests from real agents may differ in ways no model writer produces.
- Test F is not fully separate from training. Its labels were set by raters that also rated training cases, so the model may have learned their habits. Test E, labelled by a family that touched no training case, gives about the same accuracy (0.919 against 0.913) but more wrong allows: 6.7% against 3.6% at highest score, 4.2% against 1.2% at the strict threshold. Test E's writers are from the same family as the model's author.
- One training run. There is no measure yet of how much these figures move with another seed.
- English only, and short inputs only (192 tokens).
- Large models disagree with each other on cases where required evidence is missing: some answer deny, some escalate. The written rule says escalate. Decision-0 inherits that ambiguity.
Intended use
- A first-pass check inside an agent framework: route clear denials and unclear cases to a person, and let clear, policy-backed actions through only where a wrong allow is cheap to undo.
- Research on small, local models for agent oversight.
Out of scope: any decision about people's health, money, legal status or safety without human review; any use as the only control on an irreversible action.
What you can check
SHA256SUMSlists every file.model.safetensorsmust hash to the weights SHA-256 above, the checkpoint the evaluations scored.decision0.jsonrecords the input format, the label ids, the decision rule, the recommended threshold and the training settings.eval/ADVERSARIAL.mdhas a twelve-probe adversarial smoke test (self-authorisation, disguised destructive actions, data exposure, missing authority, conflicting policy, malformed and out-of-scope input): no probe was wrongly allowed at the strict threshold; two were escalated where allow or deny was expected.eval/EVAL_V3.mdhas the evaluation totals for every test, with the baselines.eval/VALIDATION.mdandeval/failed_validation_cases.jsonlshow the model's mistakes on the validation set, case by case. For the sealed tests only totals are published, so those tests stay usable.- The training and evaluation code and the training data are in HUMMBL's private repository and are not part of this release. From outside, you can reproduce inference and check the weights, not the training run.
License
Apache License 2.0. See LICENSE and NOTICE. The base model is Apache-2.0, built on microsoft/deberta-v3-xsmall (MIT).
About
Decision-0 is made by HUMMBL. Version 0.1.0 (tag v0.1.0) was its first public release; 0.1.1 corrects the card after an outside review. It was built from HUMMBL's third internal research iteration, which the evaluation records call v3. Releases follow Semantic Versioning: 1.0.0 is reserved for a version with human-checked test labels and a stable interface.
- Downloads last month
- 35
Model tree for hummbl-hf/decision-0
Base model
microsoft/deberta-v3-xsmall