Stirling-Minimo PII v0.1
0.904 average masking F2 on the PIIMB English benchmark. 140× smaller than the leaderboard leader, and the most parameter-efficient model at this accuracy.
10M parameters. ~43MB. 77 entity types. One ONNX file.
The leaderboard leader is a 1.4B-parameter mixture-of-experts model, 140× our size. We are within 1.4 points of it, and we run on a CPU or a modest previous-generation GPU, where every other option at this accuracy asks for a current-generation accelerator.
| Parameters | ~10M, 6-layer, 320-dim |
| Built from | scratch: purpose-built tokenizer and architecture, not a fine-tune of a general model |
| Model size | ~43MB, ONNX, FP32 |
| Tokenizer | byte-level BPE |
| Context window | 256 tokens |
| Labels | 77 entity types, English |
| Served system, GPU | full detection path in the Vidai engine: 1.95ms median, 7,600 requests/sec sustained on a single previous-generation GPU, ~3GB of its 24GB, so the accelerator stays free for its real workloads |
| CPU inference | 1.4ms median on a commodity 8-core CPU; 1,600 texts/sec batched on 16 cores |
Performance
PIIMB English, official piimb scorer, character-level masking F2:
| Task | F2 |
|---|---|
| gretel | 0.979 |
| nemotron-pii | 0.930 |
| ai4privacy-en | 0.917 |
| privy | 0.792 |
| Average | 0.904 |
Measured with the official scorer; independent PIIMB leaderboard validation in progress. The interesting number is the 140×, not the 1.5 points. Detection at this latency sits in the path of a live request rather than in a batch job behind it.
Masking exactness, measured on the deployed system. An external test
team, using their own corpus and their own methodology, measured the
served Vidai system at 91% boundary-exact spans; the same corpus on dslim/distilbert-NER (66M parameters), the off-the-shelf detector this system replaced in our stack, returned 19%. These figures describe the integrated system, whose span handling is part of the platform; the
published checkpoint is the masking core alone. The lesson holds either
way: over-extended spans pass containment checks and round-trip tests
while still corrupting masked text, and exact boundaries are what make
masking safe.
Why it exists
We build the Vidai Control Plane, the Sovereign layer that governs enterprise AI traffic from one boundary: what gets masked, what stays inside the network, what is recorded as evidence for an auditor.
Identification has to happen inline, on live traffic, at the point the decision is made. The control plane is engineered for efficiency, and the models are built the same way: deployment, latency and throughput are design constraints from the first day, not properties to be discovered later. Nothing available delivered all three, and that combination, rather than any one of them, was the hard problem.
Every accurate option on the market was large. The field had answered accuracy with scale, and the bill arrived in hardware and time: current-generation accelerators, roughly 150 to 200ms per inference, and throughput that collapsed under concurrency. That puts detection in a batch job behind the traffic, a report rather than a control, and it prices out the branch office, the factory floor and the customer's existing rack. Regular expressions fail the other way: fast and free, but they handle fixed forms, card numbers and emails, and cannot recognise a name they have not seen before.
We did not take that as given. We trained a 10M-parameter model that scores within 1.4 points of the leaderboard leader at 140× its size. The parameters the field spends on this task turn out not to be necessary to it.
Sovereignty, edge deployment and on-device inference are what solving that unlocked. They are consequences of the engineering, not the premise of it.
What it detects
77 entity types:
| People | person, first / last / middle name |
| Contact | email, phone, URL, IP address |
| Financial | credit card, account number |
| Identity | government ID, date of birth, biometric |
| Location | address, city, state, country |
| Credentials | keys, tokens, passwords |
| Obfuscated | deliberately disguised identifiers |
Output is typed spans with character offsets, so values can be masked, tokenised or replaced in place without string matching.
Beyond entity extraction
Classic NER answers one question: which entities are in this text. But most of what a compliance team needs to catch has no entity in it at all:
"Daniel is HIV positive." "The bailiffs have been in touch." "Marta's wages were garnished." "She's off for her scans again." "We can't make the mortgage payment this month."
Nothing in those sentences can be extracted. The sensitive fact lives in the relationship between a person and a context, often in deliberately indirect phrasing, because that is how people actually write about sensitive things. This is the problem we built the system to solve:
| Capability | What it means |
|---|---|
| Contextual detection | fires on the disclosure itself, with or without a name in the sentence |
| Euphemism handling | recognises indirect phrasings: hardship, illness, treatment, distress |
| Person-conditioned verdicts | a name plus a sensitive context is a disclosure; neither alone is |
| Category attribution | labels the disclosure: health, financial hardship, religious, political (GDPR Article 9) |
| Exact character spans | boundaries are exact, so masking and pseudonymisation are safe in place |
| Conversation awareness | in the Control Plane, a name given in one turn and a disclosure made in another is still caught |
That turns a detection log into an audit trail: the difference between redacting a document and showing a regulator what was disclosed, when, and under which article.
This release publishes the model, not just the entity core.
onnx/model.onnx ships the entity head the leaderboard measures;
onnx/model-full.onnx adds the person-name specialist and the
contextual-sensitivity head in one multi-output build (see Using it).
The fuller capability above, person-conditioned verdicts, category
attribution and the audit record, runs in the Vidai Control Plane, in
the same efficiency envelope, combining these heads with gating and
fusion. For production deployment, we can help you configure the right
combination of detection, gating and audit for your use case.
Built for GDPR
Most detection tools stop at "found personal data". GDPR asks three harder questions: what category was disclosed, did it leave your control, and can you prove what happened. This system answers all three.
Article 9, attributed. Special-category data is detected and labelled, not just flagged:
| GDPR Article 9 category | How it is caught |
|---|---|
| Health | clinical, euphemistic and contextual health disclosure |
| Religious / philosophical belief | observance, practice, affiliation |
| Political opinions | party, voting, campaigning |
| Trade union membership | membership and activity |
| Sexual orientation | disclosure |
| Genetic data | tests, markers, heritage |
| Biometric data | identifiers and scans |
Plus the Article 6 basics: names, contact details, financial data, identifiers, locations, across 77 entity types. Attribution is what turns a detection log into evidence: the difference between "something was redacted" and "a health disclosure about a named individual was blocked at the boundary, 14:02, Article 9".
Chapter V, by not transferring. Checking text with a hosted detection API is itself a processing of personal data by a third party: a processor on the register, a DPA, a transfer mechanism. A 43MB file that runs where the data already is needs none of that. It removes an entire chapter of compliance work rather than documenting it.
Article 32, technical measures. Inline masking and pseudonymisation before third-party model calls, at request speed, are exactly the technical measures Article 32 asks for, and pseudonymisation is named in the article itself.
Article 5(2), provable. The Control Plane records the audit trail: what was disclosed, when, which category, under which article. Records of processing and DPIA evidence come from the same stream.
Compliance itself remains the controller's responsibility. This is the technical layer that makes it achievable without a compliance team reading every transcript.
Where it runs
At 43MB with no accelerator required, the constraint that usually decides where a detector can live stops applying.
Sovereign by construction. A file you deploy, not an API you call. No outbound request, no third-party processor on the register, nothing to disclose in a DPA, and air-gap ready: it runs with no network at all. For regulated organisations that is not a preference; it is the only shape a detector is allowed to take.
At the edge. Identification where data is created rather than at a central point it must reach first: branch systems, field devices, regional deployments.
On a handset. The envelope fits on-device inference: a personal data firewall, where the phone decides what leaves it and the decision never leaves either.
The corporate version of that is the whole argument. Checking text with a hosted API means transmitting it, customer records, support transcripts, internal correspondence, to a third party. A 43MB file running locally means the material never moves.
Using it
Load the ONNX graph with onnxruntime. Predictions are argmax per token, with same-type runs grouped into spans: the standard "simple" aggregation. It ships fully calibrated, with the operating point included in the exported weights; there is nothing to configure and no threshold to tune.
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
REPO_ID = "VidaiUK/stirling-minimo-pii-10m"
# 1. Load tokenizer and fetch the ONNX graph directly from the Hub
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/model.onnx")
session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
text = "Daniel's email is daniel@example.com"
# 2. Tokenize with character offset mappings (byte-level BPE)
inputs = tokenizer(text, return_tensors="np", return_offsets_mapping=True)
offsets = inputs.pop("offset_mapping")[0]
# 3. Inference
ort_inputs = {k: v for k, v in inputs.items()
if k in ("input_ids", "attention_mask")}
logits = session.run(None, ort_inputs)[0] # Shape: (batch, seq, 78)
# 4. Spans (the optimal F2 operating point is baked into the head bias)
predictions = logits[0].argmax(axis=-1)
# Pair predictions with offsets[i] to extract exact [start:end] slices
ONNX is the native format (the architecture is custom, so use onnxruntime rather than a transformers model class). The full label set:
77 entity types (click to expand)
account_number, accountname, accountnumber, address, age, amount, bank_routing_number, bic, biometric, biometric_identifier, bitcoinaddress, buildingnumber, certificate_license_number, city, companyname, coordinate, country, credential, credit_card, credit_debit_card, creditcardcvv, creditcardissuer, currency, currencycode, currencyname, currencysymbol, customer_id, date, date_of_birth, device_identifier, dob, education_level, email, employee_id, employment_status, ethereumaddress, fax_number, firstname, gender, government_id, health_plan_beneficiary_number, http_cookie, ip_address, jobtype, language, lastname, litecoinaddress, mac, mac_address, maskednumber, medical_id, middlename, nearbygpscoordinate, obfuscated_pii, ordinaldirection, organization, person, personal_attribute, phone, political_view, postcode, prefix, race_ethnicity, religious_belief, secondaryaddress, sexuality, state, swift_bic, time, title, unique_id, unique_identifier, url, useragent, vehicle, vehicle_identifier, vehiclevrm
The other heads: onnx/model-full.onnx
The leaderboard figure measures one output: the entity head, which is
what onnx/model.onnx ships. The same encoder carries two further
heads, published together in onnx/model-full.onnx. One inference pass
returns all three:
| Output | Shape | What it is |
|---|---|---|
logits |
(batch, seq, 78) | the entity head; identical predictions to onnx/model.onnx |
name_logits |
(batch, seq, 1) | person-name specialist, per token |
ctx_logits |
(batch, 2) | contextual-sensitivity verdict, per document |
The entity output is the benchmark configuration, unchanged: the same calibration, the same predictions. The name head is a higher-sensitivity person signal, useful where name coverage matters beyond what the entity head alone provides. The context verdict fires on disclosure phrasing, a person plus a sensitive context, rather than on any entity: the cases classic NER cannot see.
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
REPO_ID = "VidaiUK/stirling-minimo-pii-10m"
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/model-full.onnx")
session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
def analyse(text):
inputs = tokenizer(text, return_tensors="np", return_offsets_mapping=True)
offsets = inputs.pop("offset_mapping")[0]
ort_inputs = {k: v for k, v in inputs.items()
if k in ("input_ids", "attention_mask")}
logits, name_logits, ctx_logits = session.run(None, ort_inputs)
predictions = logits[0].argmax(axis=-1) # entity head, per token
name_probs = 1 / (1 + np.exp(-name_logits[0].squeeze(-1))) # person specialist
ctx_prob = float(np.exp(ctx_logits[0][1]) / np.exp(ctx_logits[0]).sum())
spans = [] # group same-type runs into character spans
current = None
for i, (start, end) in enumerate(offsets):
if predictions[i] == 0:
if current: spans.append(current); current = None
continue
if current and current[2] == predictions[i]:
current = (current[0], int(end), predictions[i])
else:
if current: spans.append(current)
current = (int(start), int(end), predictions[i])
if current: spans.append(current)
for s, e, _ in spans:
print(f"entity: {text[s:e]!r}")
for i, p in enumerate(name_probs):
if p > 0.8:
s, e = offsets[i]
print(f"name p={p:.2f}: {text[s:e]!r}")
print(f"context sensitivity: {ctx_prob:.2f}")
analyse("Sarah Whitfield called about the insurance claim")
Output:
entity: 'Sarah Whitfield'
name p=1.00: 'Sarah'
name p=1.00: 'Wh'
name p=1.00: 'it'
name p=1.00: 'field'
context sensitivity: 1.00
The token pieces in the name readout are sub-word units of the same name. The entity head is the primary interface and the one the benchmark measures; the additional heads are complementary signals. In the served Vidai system they are combined with gating and span handling; consumed standalone they are confidence signals rather than finished detections.
If you mask rather than only detect, score boundary equality, not containment: an over-extended span passes both a containment check and a replace-then-restore round-trip, and still corrupts the text.
Limitations
English only. A multilingual model is in development.
Names. Most names, across a wide range of origins, are detected. A small number of uncommon names, including some written with accents or special characters, may be fragmented or missed. The multi-output build's name specialist (see Using it) extends coverage on exactly these cases. Misses are deterministic, so the same name fails the same way every time and surfaces in testing rather than as random leakage. Where coverage must be guaranteed, pair with a known-names list.
Fairness is measured, not assumed. In the integrated system, name detection is evaluated per group across 19 ethnic origins, and per-group figures are available on request. Equal protection across origins is a design requirement with a number attached to it, which is rare in this category.
Context window. This release processes up to 256 tokens per pass. For longer texts, chunk with a 32-token overlap stride and merge adjacent identical spans. Long-document handling is provided in the Vidai Control Plane; we can advise on chunking strategy for your document types.
Recall-biased. On text where PII is sparse, this core fires more than a precision-tuned detector would. Production deployments should tune the operating point, gating and span handling to their traffic profile; the Vidai Control Plane provides this, and we can advise on the right setup for yours.
The rest of the series
Specialist models stack. Same frozen backbone, so you run several at once rather than choosing between them:
| Stirling-Minimo model | Detects |
|---|---|
| Health | clinical and euphemistic health disclosure |
| Financial / fintech | hardship, collections, insolvency, weak-signal financial language |
| Compensation | salary, pay bands, offer and remuneration data |
| Prompt injection | injection and jailbreak attempts |
| Custom | your own organisation's sensitive-data categories |
"Sensitive" is rarely just the regulated categories. It is the deal codenames, the internal project names, the pre-announcement figures. No general model knows them, which is why the last row is the one most organisations need.
Specialising costs another file, not a bigger one. Four models at ~43MB each, core, health, compensation and injection, run together in less memory than one mid-size model needs for its weights alone. A stack of small specialists is cheaper than one general model, and each is better at its own job.
Every model keeps the same deployment shape: a file you run where the data is. Together they are the detection layer of the Vidai Control Plane, which runs them inline and distributed, and handles policy, routing, masking, pseudonymisation and the audit record. Licensing and access: vidai.uk
Intended use
Built for. Identifying personal data leaving an organisation through AI systems; masking or pseudonymising before third-party model calls; Sovereign, on-prem, edge and on-device filtering where text must not be transmitted to be checked.
Not for. Profiling or discriminating on the basis of detected data. Legal determinations about data classification. Sole-control use where a missed detection causes serious harm: put deterministic rules alongside.
Training
Trained from scratch. The tokenizer and the architecture are ours: a byte-level BPE vocabulary built for this task, and a compact transformer designed for the deployment envelope rather than inherited from it. This is not a fine-tuned general-purpose model carrying another model's size, vocabulary and provenance into yours.
Pretraining is masked-language modelling over publicly available text: Wikipedia and the Enron corpus of real workplace email, followed by multi-task fine-tuning over public PII corpora and a substantial body of purpose-generated synthetic data.
Synthetic data, built from regulatory experience. The corpus was not assembled; it was authored. It comes out of years of practical work in data privacy and regulation, which teaches you what sensitive data actually looks like in flight: the indirect phrasings people use for illness and hardship, the disclosures that never carry an entity, the formats leaks take, the categories a regulator will ask about. None of that is in public corpora, because nobody publishes their sensitive incidents. The synthetic side of the corpus encodes exactly that experience: euphemistic and indirect disclosure, diverse personal names across many origins, disguised identifiers and adversarial formats. The generation methodology is proprietary; the expertise shows up in the results.
With thanks to:
| Source | Licence |
|---|---|
| Ai4Privacy OpenPII (train split) | CC BY 4.0 |
| NVIDIA Nemotron-PII (train split) | CC BY 4.0 |
Also used: public-domain US government name lists and in-house synthetic and template-generated data.
Test hygiene. All public training sources were audited against the PIIMB English test set: zero test documents in any training file, by exact-match and n-gram audit. Nemotron-PII is template-generated, so phrase overlap between its own splits is inherent to the source.
Release status
v0.1, public preview. First of the Stirling-Minimo series. Benchmark results are from the official scorer on the public test set; independent PIIMB leaderboard validation is in progress. The checkpoint is complete; what is provisional is its shape: weights, label set, context window and output format may change before v1.0. Pin the revision if you build against it. Failure cases on real data are why this is public.
Licence
CC BY-NC 4.0, free for research, evaluation, personal and non-commercial use. Commercial use requires a separate licence: vidai.uk
- Downloads last month
- 23