Ariadne Laya Privacy
Returns six privacy flags: email, telephone, credit card, street address, date of birth and IPv4 address. This 4.2 MB specialist interface runs on a shared frozen Laya base. Switching between compatible Ariadne specialists replaces about 1 million parameters, while the 421 million parameter base stays in memory. Each specialist uses its own small interface.
Use
Install the included Python wheel from this downloaded model folder:
pip install ./ariadne_specialists-0.2.0a1-py3-none-any.whl
import ariadne
from ariadne.specialists import Privacy
model = ariadne.load_specialist(Privacy, model=".")
result = model('Please contact Alex at alex@example.com.')
print({flag: prediction.label for flag, prediction in result.items()})
The base downloads automatically and is cached. Choose a device with device="cpu" or device="cuda:1". A list of inputs returns a list of results. Scores have not been recalibrated for this task.
Load from Hugging Face
After installing the included wheel, you can load this repository directly:
import ariadne
from ariadne.specialists import Privacy
model = ariadne.load_specialist(Privacy, model="GoatHerder/Ariadne-Laya-PrivacyFlags")
Use the explicit model= argument with this preview wheel. The interface and pinned base are downloaded automatically and cached. Pass revision="<commit hash>" to pin a particular interface version.
Interface
The interface is a 1,024 × 1,024 linear projection plus a 1,024-element bias: 1,049,600 trainable parameters. It sits after the base's native embeddings and before encoder block 0. It starts as the identity; training updates only this projection. The shared base has 421,293,827 parameters. Compatible specialists share one resident base in the same Python process and on the same device.
Results
Local evaluation uses the same 1,000 inputs for every model. Accuracy is per decision; macro-F1 averages the task's classes (and flag namespaces for Privacy).
| Model | Accuracy | Macro-F1 |
|---|---|---|
| Base Laya | 76.05% | 69.75% |
| Ariadne Privacy | 97.22% | 95.21% |
| Horizon-Labs/pii-redactor-small | 97.63% | 96.52% |
| TF-IDF + logistic regression | 94.90% | 90.03% |
Privacy positive-class macro-F1: base 56.19%, interface 92.25%, public comparator 94.62%. Per-flag precision and recall are in metrics.json.
Fine-tuned on Gretel plus other PII sources. BIO token predictions are aggregated into six document-presence flags, using windows over the full text. IPv4 maps to the broader IP_ADDRESS tag. Training-row overlap is unverified. This table does not establish a common unseen-test ranking. metrics.json records model revisions, comparison methods, per-class results and existing-task retention.
Training and scope
Trained on gretelai/gretel-pii-masking-en-v1, revision e06eb1499ca8d54470f085021cd8e54f9efac7fd. Prepared train/validation/test sizes: 3,000 / 500 / 1,000. Overlength exclusions: {'train': 0, 'validation': 0, 'test': 0}.
One seed (0); epoch 2 selected by validation loss, training stopped after epoch 10. LR 1e-4, minimum 10 epochs, patience 3. Only the 1,049,600 exact-identity-initialized interface parameters were trained. The original embeddings, 28 encoder blocks and decision heads stayed frozen and in evaluation mode. Deterministic GPU settings were enabled.
- Six document-level type-presence flags from synthetic annotated text. This model does not locate or redact spans.
- All source documents contain some PII; absence is measured per type, not on wholly PII-free documents. Stable samples of 3,000/500/1,000 documents.
English only. Unsupported or ambiguous inputs still receive a prediction. Source datasets retain their own licences. This checkpoint is one training run; it does not establish across-seed variance.
Base revision: 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Independent adaptation; no affiliation with the original Laya authors.
Model tree for GoatHerder/Ariadne-Laya-PrivacyFlags
Base model
convaiinnovations/laya