Phonon-2 for phonon-2-c
Phonon-2, Fermion Research's five-value re-training of NVIDIA's
Parakeet-TDT 0.6B v3, repacked for phonon-2-c, a pure-C inference
engine for CPUs with no dependencies beyond libc and pthreads. These are the weights of the Phonon-2 release
re-encoded, not re-quantised: the encoder's five-value linears stay five-value, and the engine decodes identical
token ids to the transformers reference on the same features.
| File | Size | sha256 |
|---|---|---|
phonon-2-c-weights-v1.tar.gz |
187 MB | 48137a58011c4b0bc0a97087ea7e868daabf1525ca8447414be230b21faf99e6 |
It unpacks to c_weights/ (255 MB): weights.bin, weights.json (the tensor index) and tokens.txt (the
vocabulary). The format is phonon-2-c-weights-v1; its layout is described at the top of convert_weights.py.
Run it
git clone https://github.com/eschmidbauer/phonon-2-c && cd phonon-2-c
make -C csrc
curl -L https://huggingface.co/eschmidbauer/phonon-2-c/resolve/main/phonon-2-c-weights-v1.tar.gz | tar xz
./csrc/phonon2 c_weights recording.wav
./csrc/phonon2 --json c_weights recording.wav # word timestamps
scripts/get_weights.sh in the engine repository (also make -C csrc weights) is the same download with the
archive's sha256 checked against the value pinned there. Any WAV works (PCM or float, any sample rate, mono or
multi-channel), and --stream reads raw 16-bit PCM from stdin. The engine's README covers the command line, the C library and the Python bindings. On an Apple M5 Max the
default mode transcribes a 256 s call in 1.7 s and a 3 s utterance in 0.03 s.
Accuracy
Measured with the engine on LibriSpeech, scored with the Whisper English normaliser (what the Open ASR Leaderboard
uses). exact expands the five values to fp32; int8w, the default, runs int8 dot products over the same codes.
| Set | int8w (default) |
exact |
|---|---|---|
| test-clean (2,620 utterances) | 2.16 % | 2.16 % |
| test-other (2,939 utterances) | 4.34 % | 4.33 % |
Long recordings are decoded in pause-aligned 25–35 s windows, the way Phonon-2's own engine does it; on
concatenated LibriSpeech chapters that scores 1.79 % (test-clean) and 3.19 % (test-other). The engine's README has
the protocol, the parity test against transformers and the speed table.
What is inside
| Tensors | Count | Stored as |
|---|---|---|
| Encoder linears (feed-forward, fused Q/K/V, position, output, pointwise convs) | 216 | five-value: 2-bit signs + 1-bit lo/hi flags, [in, out], fp32 lo/hi per output, 229 MB |
| Subsampling linear, encoder projector | 2 | int8 with one fp32 scale per output column, exact from the container's int6 |
| Embedding, LSTM, decoder projector, joint head | 7 | int8 with one fp32 scale per row, exact, 17 MB |
| Norms, biases, subsampling convs, depthwise conv with its BatchNorm folded in | 354 | fp32 |
Everything is exact except the BatchNorm fold, which is computed in double precision and stored as fp32.
Verifying and rebuilding
shasum -a 256 -c --ignore-missing SHA256SUMS.txt # the archive, and the three files once unpacked
build.sh rebuilds the archive from Phonon-2's release and checks the result against those sums. It downloads
phonon-2.bps.tar.zst (164 MB, sha256 98125795b6dda72f5c6eee9ba33d19815df65dcb18b50a357bf9f73c9935309e) from
FermionResearch/Phonon-2, checks the container inside it (model.fermion, sha256
4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468, also recorded in weights.json under _meta),
converts it with convert_weights.py and packs c_weights/ deterministically (ustar, fixed order, zero timestamps
and owners, no gzip header name), so a rebuild gives this exact file:
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt # numpy, zstandard, huggingface_hub
./build.sh
convert_weights.py is a verbatim copy of the engine's converter, and fermion_container.py is Fermion Research's
reference reader for the container, unchanged from the Phonon-2 repository.
Licence
The weights are CC-BY-4.0 (LICENSE-WEIGHTS-CC-BY-4.0.txt), the licence of Phonon-2 and of NVIDIA's Parakeet-TDT
0.6B v3 they derive from; the attribution chain, the changes made and the training-data terms are in
NOTICE. convert_weights.py and build.sh are MIT (LICENSE-CODE-MIT.txt), as is the engine;
fermion_container.py is Apache-2.0 (LICENSE-CODE-Apache-2.0.txt), Fermion Research.