Statim Decide Multilingual Base: PII adapter
Licence notice (2026-10-03). The base model this adapter requires (multilingual 0.7.0) was trained partly on data that is non-commercial, ShareAlike or under an unknown licence. From 2026-10-03 the base and this adapter are offered only under PolyForm Noncommercial 1.0.0; the Small Business and Free Trial licences and the commercial licence do not apply to them. A replacement trained only on cleared data will follow. Details: DATA_LICENSES.md.
A LoRA adapter that improves the PII decisions of
Beko2210/statim-decide-multilingual-base 0.7.0.
Checked with Statim 0.8.3: the published f32 file loads it merged at load, the q8_0 file as runtime LoRA. LoRA adapters need Statim 0.8.0 or later.
Quick start
hf download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
hf download Beko2210/statim-decide-multilingual-base-pii statim-decide-multilingual-base-pii.lora.gguf --local-dir models
statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --adapter multilingual:pii=models/statim-decide-multilingual-base-pii.lora.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{"state": {"text": "Hi, this is Anna Schmidt. Call me back at +49 170 1234567."}, "questions": {"pii": {"type": "noul", "instructions": "Does the text contain a phone number?"}}, "adapter": "pii"}'
curl -s localhost:8080/v1/systemone -d '{"state": {"text": "Hi, this is Anna Schmidt. Call me back at +49 170 1234567."}, "questions": {"pii": {"type": "noul", "instructions": "Does the text contain a phone number?"}}, "adapter": "auto"}'
Python (Python SDK 0.8.3 or later):
from statim import Client
client = Client("http://127.0.0.1:8080")
client.decide({"text": "Hi, this is Anna Schmidt. Call me back at +49 170 1234567."}, {"pii": {"type": "noul", "instructions": "Does the text contain a phone number?"}}, adapter="pii")
Results (experiment's f32 run)
Each cell uses 150 items (seed 20260927).
| Language | Base | Adapter | Change (points) | Verdict | Qwen3-8B zero-shot |
|---|---|---|---|---|---|
| ar | 0.8467 | 0.9067 | +6.00 | within noise | 0.893 |
| de | 0.8400 | 0.8667 | +2.67 | within noise | 0.893 |
| en | 0.8933 | 0.9267 | +3.34 | within noise | 0.887 |
| es | 0.8533 | 0.8933 | +4.00 | within noise | 0.867 |
| fr | 0.8933 | 0.9400 | +4.67 | within noise | 0.880 |
| it | 0.8400 | 0.9200 | +8.00 | gain (2 SE) | 0.873 |
| ja | 0.8467 | 0.9267 | +8.00 | gain (2 SE) | 0.893 |
| nl | 0.7600 | 0.8667 | +10.67 | gain (2 SE) | 0.800 |
| ru | 0.9067 | 0.9467 | +4.00 | within noise | 0.947 |
| sv | 0.8333 | 0.8733 | +4.00 | within noise | 0.800 |
| zh | 0.9000 | 0.9467 | +4.67 | within noise | 0.920 |
| Mean | 0.8558 | 0.9103 | +5.46 | — | 0.878 |
The pooled family change is +5.46 points with 2 SE = 2.22 points (gain).
A single cell of 150 items rarely clears 2 SE on its own; the decision uses the pooled family.
Decision rule: promote when the category family gains more than 2 standard errors, pooled by rows or by suites, and nothing regresses. A regression is a pooled drop beyond 2 standard errors under either pooling, or a drop in one language cell that stays significant after Holm-Bonferroni (family-wise 5 %).
Statim is trained on this category, while Qwen3-8B runs zero-shot.
Checked on the published files
| Weights | Adapter mode | Base mean | Adapter mean | Change (points) | Cell for cell as in the experiment |
|---|---|---|---|---|---|
| f32 | merged at load | 0.8558 | 0.9103 | +5.46 | yes |
| q8_0 | runtime LoRA | 0.8558 | 0.9109 | +5.52 | — (experiment: f32) |
Training
| Source | Rows | Licence |
|---|---|---|
E3-JSI/synthetic-multi-pii-ner-v1 (default) |
2,971 | MIT |
Wismut/nym-pii-multilingual-data (default) |
6,200 | MIT |
gretelai/gretel-pii-masking-en-v1 (default) |
6,200 | Apache-2.0 |
gretelai/synthetic_pii_finance_multilingual (default) |
6,200 | Apache-2.0 |
nvidia/Nemotron-PII (default) |
6,200 | CC-BY-4.0 |
urchade/synthetic-pii-ner-mistral-v1 (data.json) |
6,200 | Apache-2.0 |
LoRA rank 16, alpha 32.0, dropout 0.05;
target modules Wqkv, Wo, Wi; 88 wrapped modules and
3,379,200 trainable parameters.
- Items: 33,571 train, 400 dev.
- Updates: 1,172.
- Dev accuracy: 0.8375 before, 0.8725 after.
- Time: 1631.7 seconds; peak memory: 2,332 MB.
Provenance
- Adapter GGUF SHA-256:
2991a33b5d9db4f5b679081885faab1f84289d6eae7660830447e790bd168b29 - PEFT safetensors SHA-256:
f8e7e674ab868b312ff32f403552daa436e7437411e83fa0ece60187431a2d6f - Training checkpoint
model.safetensorsSHA-256:c44425f14ac9d55508f73a6f371e4e2802ed59e646287c3abbec10b653a19840 - Base fingerprint:
e7a8fa743b4920850167b587226e339d651be0244fcc6a9e24fff8ed5073bae5 - Training mixture SHA-256:
b8a87e85f3bf72b61509555aba9148358760d3329ecf37a36cebd94f9358cb03 - Source registry SHA-256:
309ce8ec262886b3bfaa529934151ae0c1879595c26301e4269da276d95cbfcd
Experiment commands (paths relative to the Statim repository):
train:'.venv-train/bin/python' 'tools/finetune/train_lora.py' 'models/laya-multilingual-v9' --mixture 'data/mixture-v8.jsonl.gz' --category pii --registry 'tools/finetune/sources/v6-keep.json' --out 'models/lora-exp1/pii' --device cuda --epochs 2convert:'.venv/bin/python' 'tools/convert_lora.py' 'models/lora-exp1/pii' -o 'models/lora-exp1/pii.lora.gguf' --base 'models/laya-multilingual-v9-f32.gguf' --category pii --name piiserve:'build-vk/statim' serve -m 'multilingual=models/laya-multilingual-v9-f32.gguf' --adapter 'multilingual:pii=models/lora-exp1/pii.lora.gguf' --device vulkan --threads 16 --port 8098 --no-access-log --inference-timeout 600eval_base:'.venv/bin/python' 'bench/eval_categories.py' --strict --url http://127.0.0.1:8098 --model multilingual --suites pii --n 150 --seed 20260927 --exclude-mixture 'data/mixture-v8.jsonl.gz' --out 'models/lora-exp1/pii.base.jsonl'eval_adapter:'.venv/bin/python' 'bench/eval_categories.py' --strict --url http://127.0.0.1:8098 --model multilingual --suites pii --n 150 --seed 20260927 --exclude-mixture 'data/mixture-v8.jsonl.gz' --adapter pii --out 'models/lora-exp1/pii.adapter.jsonl'
Protocol and experiment: docs/ADAPTERS.md. Full reproduction instructions: REPRODUCE.md.
Intended use and limits
- This adapter only helps its category; route requests with
"pii"or"auto". - Bound to
Beko2210/statim-decide-multilingual-base0.7.0 by the base fingerprinte7a8fa743b492085…(statim.lora.base_fingerprint, SHA-256 over the checkpoint's norm and bias tensors); Statim refuses the adapter on a base whose fingerprint differs. - Languages outside the evaluated list are untested.
- Each language cell has 150 items.
- Do not automate decisions about people without human review.
Licence
The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.
Training data attribution is listed source by source above, with the row count and licence read from the experiment registry.
- Downloads last month
- 13
We're not able to determine the quantization variants.
Model tree for Beko2210/statim-decide-multilingual-base-pii
Base model
convaiinnovations/laya-multilingualEvaluation results
- accuracy on pii ar (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.907
- accuracy on pii de (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.867
- accuracy on pii en (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.927
- accuracy on pii es (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.893
- accuracy on pii fr (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.940
- accuracy on pii it (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.920
- accuracy on pii ja (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.927
- accuracy on pii nl (n=150, seed=20260927, z=2.0, alpha=0.05)self-reported0.867