Mark-38M

A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (ak), Arabic (ar), Azerbaijani (az), Catalan (ca), Czech (cs), Welsh (cy), Ewe (ee), Spanish (es), Pulaar (ff), French (fr), Irish (ga), Guaraní (gn), Hausa (ha), Hebrew (he), Croatian (hr), Haitian Creole (ht), Hungarian (hu), Igbo (ig), Kurdish (ku), Lingala (ln), Lithuanian (lt), Latvian (lv), Māori (mi), Polish (pl), Portuguese (pt), Quechua (qu), Romanian (ro), Slovak (sk), Slovenian (sl), Samoan (sm), Serbian (sr), Turkmen (tk), Turkish (tr), Uzbek (uz), Vietnamese (vi), Wolof (wo), and Yorùbá (yo).

On the official academic held-out Yorùbá YAD test set (3,330 sentences, 142k characters), it achieves a 15.88% Diacritic Error Rate (DER), 19.38% Word Error Rate (WER), and 5.58% Character Error Rate (CER) with 0.0139% text corruption (zero invented or dropped words), improving over classical baselines by 53.22 points.

Across the 37-language joint evaluation suite, it achieves 93.69% macro marked-position accuracy with a composite score of 0.8419. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes.

The model is exported into native on-device formats: a 41.75 MB INT8 ONNX graph for CPU and WebAssembly, and a compiled Core ML package for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with 99.40% character parity across 828 evaluation characters.


The Task

Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints:

  1. Tone Languages (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like ba can represent bá (to accompany/meet), bà (to perch/alight), or ba (to hide). Missing tones invert negation, tense, and aspect.
  2. Abjads (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (ḥarakāt: fatḥah, ḍammah, kasrah, sukūn, shaddah) or case nunation (tanwīn). Syntax (i'rab), voice (active vs passive), and word semantics require contextual disambiguation across the sentence.
  3. Latin Orthographies (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (comio vs comió), nominal cases, and distinct phonemes (c vs ç, s vs ş, a vs ă).

Why Bytes, Not Subwords or Characters

Subword tokenizers (BPE, WordPiece) break down on diacritized text:

  • Unmarked input and marked output share base characters but differ in byte sequences.
  • Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably.
  • Subword vocabularies cannot generalize to rare combining tone stacks (e.g. M:0300+0301).

Mark operates directly on the 256 UTF-8 byte values. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary:

  • KEEP: Maintain the underlying byte.
  • P:src->dst: Substitute single character (e.g. P:a->á).
  • M:hex: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (NFC).

Invariant Guarantee: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words.


Results

Official Yorùbá YAD Test Benchmark (3,330 Sentences)

Evaluated on the academic gold-standard Yorùbá YAD test corpus (Asahiah et al., Orife):

System DER (Diacritic Error) WER (Word Error) CER (Char Error) Hallucination / Invention Inference Latency
Mark-38M 15.88% 19.38% 5.58% 0.0139% 50.07 ms / sent
Unmarked Input (Identity Baseline) 69.10% 78.40% 21.20% 0.0000% —
Absolute Improvement -53.22% -59.02% -15.62% Zero text drift 20 sent/sec (Apple Silicon)

37-Language Joint Accuracy Report

Held-out evaluation report across all 37 languages:

Code Language Marked-Position Accuracy Code Language Marked-Position Accuracy
az Azerbaijani 99.65% ee Ewe 97.39%
fr French 99.47% sl Slovenian 97.39%
tr Turkish 99.02% sk Slovak 97.21%
gn Guaraní 88.61% lv Latvian 96.86%
es Spanish 98.22% ku Kurdish (Kurmanji) 96.51%
pt Portuguese 98.09% ig Igbo 96.30%
ga Irish 98.07% hu Hungarian 95.60%
sr Serbian 98.07% cs Czech 95.47%
pl Polish 97.89% ht Haitian Creole 95.47%
ak Akan (Twi) 97.88% ar Arabic 94.86%
hr Croatian 97.84% ca Catalan 94.85%
tk Turkmen 97.80% wo Wolof 93.07%
lt Lithuanian 97.79% ha Hausa 92.18%
ro Romanian 97.41% sm Samoan 90.98%
vi Vietnamese 90.98% qu Quechua 89.78%
he Hebrew 89.40% uz Uzbek 89.15%
yo Yorùbá 87.39% ff Pulaar (Fula) 85.56%
mi Māori 81.15% ln Lingala 78.56%
cy Welsh 74.68%
  • Macro marked-position accuracy: 93.69%
  • Worst-case language: Welsh (cy) at 74.68%
  • Composite score: 0.8419

Quantization & Edge Footprint

Format Precision File Size Recommended Target Latency (CPU / Apple NE)
mark_int8.onnx Dynamic INT8 41.75 MB Edge CPU, Mobile, Browser (WASM) 71.49 ms (4-thread CPU)
mark_fp16.onnx Float16 80.01 MB Mobile GPUs, WebGPU 25.10 ms (GPU)
mark.mlpackage 8-bit Core ML 42.10 MB Apple Neural Engine (iOS, macOS) 12.40 ms (ANE)
mark_fp32.onnx Float32 158.81 MB Reference server baseline 135.15 ms (1-thread CPU)

Install & Integration

Python (ONNX Runtime)

from mark import restore

# Standard inference
restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo")
# -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."

# Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4)
stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4)
# -> "El niño comió jamón en la mañana."

iOS / macOS (Swift Package Manager)

Add to your Package.swift:

dependencies: [
    .package(url: "https://huggingface.co/Mythologic/mark", branch: "main")
]

Declare dependency in your target:

.product(name: "Mark", package: "mark")

Direct inference on Apple Neural Engine (Swift 6.3):

import Mark

let mark = try Mark()

// Yorùbá
let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo")
// -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."

// Arabic
let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar")
// -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ"

// Spanish (with calibrated margin: 0.4 for zero typing jitter)
let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4)
// -> "El niño comió jamón en la mañana."

Browser & Node.js (npm)

Install via npm:

npm i @mythologic/mark

Inference via ONNX Runtime Web (WASM / WebGPU):

import { Mark, restore } from "@mythologic/mark";

// Direct helper
const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo");
// -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé.

const arabic = await restore("ذهب الولد الى المدرسة", "ar");
// -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ

// Reusable instance with calibrated anti-flicker margin
const mark = new Mark();
await mark.init();
const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 });
// -> El niño comió jamón en la mañana.

Architecture

Attribute Specification
Parameters, total 38,735,217 (38.7M)
Encoder layers 16 Transformer encoder layers
Hidden dimension 384
Intermediate dimension 1,536 (SwiGLU feed-forward)
Attention heads 8 heads (head dimension 48)
Positions Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form)
Convolution stem 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192)
Vocabulary 257 (256 raw UTF-8 bytes + language conditioning)
Classifier head Linear projection to 1,073 tag classes
Context window 512 bytes with sentence/whitespace sliding window

Training Setup

  • Dataset volume: ~786M tokens across 37 languages.
  • Loss formulation: Class-Balanced focal loss (cb, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0).
  • Sampling: Temperature-scaled power law ($\tau = 0.5$) over language shards.
  • Schedule: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence).
  • Optimizer: Muon (matrix trunk) + AdamW (embeddings & classifier).
  • Hardware: Dedicated high-throughput accelerator cluster.

Limitations & Failure Modes

  1. Zero-Context Homographs: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish si [if] vs sí [yes], Arabic qtr -> qaṭara vs qaṭr) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency.
  2. Tail Language Shard Volume: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed 95–99% accuracy. Low-resource tail languages in the joint mixture (Welsh cy at 74.68%, Lingala ln at 78.56%) exhibit lower precision due to smaller corpus volume.
  3. Dialectal Writing: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy.
  4. Window Chunk Boundaries: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss.

Files

File Format Size Description
mark_int8.onnx ONNX (INT8) 41.75 MB Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly
mark_fp16.onnx ONNX (FP16) 80.01 MB Half-precision graph for GPUs and Neural Engines
mark_fp32.onnx ONNX (FP32) 158.81 MB Full-precision reference model
mark.mlpackage.zip Core ML 36.16 MB Compiled Core ML package for Apple Neural Engine
config.json JSON 1 KB Model architectural hyperparameters
tags.json JSON 40 KB 1,073 tag operation mappings
languages.json JSON 1 KB 37 ISO 639-1 language codes

Author & Citation

Developed by Ainouche Abderahmane and Mythologic.

Released under the Apache 2.0 License.

@software{ainouche_mythologic_mark_2026,
  title  = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages},
  author = {Ainouche, Abderahmane and Mythologic},
  year   = {2026},
  url    = {https://huggingface.co/mythologic/mark},
  note   = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages}
}

Part of Mythologic.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Mythologic/mark