Phonon-2 for phonon-2-c

Phonon-2, Fermion Research's five-value re-training of NVIDIA's Parakeet-TDT 0.6B v3, repacked for phonon-2-c, a pure-C inference engine for CPUs with no dependencies beyond libc and pthreads. These are the weights of the Phonon-2 release re-encoded, not re-quantised: the encoder's five-value linears stay five-value, and the engine decodes identical token ids to the transformers reference on the same features.

File Size sha256
phonon-2-c-weights-v1.tar.gz 187 MB 48137a58011c4b0bc0a97087ea7e868daabf1525ca8447414be230b21faf99e6

It unpacks to c_weights/ (255 MB): weights.bin, weights.json (the tensor index) and tokens.txt (the vocabulary). The format is phonon-2-c-weights-v1; its layout is described at the top of convert_weights.py.

Run it

git clone https://github.com/eschmidbauer/phonon-2-c && cd phonon-2-c
make -C csrc
curl -L https://huggingface.co/eschmidbauer/phonon-2-c/resolve/main/phonon-2-c-weights-v1.tar.gz | tar xz
./csrc/phonon2 c_weights recording.wav
./csrc/phonon2 --json c_weights recording.wav     # word timestamps

scripts/get_weights.sh in the engine repository (also make -C csrc weights) is the same download with the archive's sha256 checked against the value pinned there. Any WAV works (PCM or float, any sample rate, mono or multi-channel), and --stream reads raw 16-bit PCM from stdin. The engine's README covers the command line, the C library and the Python bindings. On an Apple M5 Max the default mode transcribes a 256 s call in 1.7 s and a 3 s utterance in 0.03 s.

Accuracy

Measured with the engine on LibriSpeech, scored with the Whisper English normaliser (what the Open ASR Leaderboard uses). exact expands the five values to fp32; int8w, the default, runs int8 dot products over the same codes.

Set int8w (default) exact
test-clean (2,620 utterances) 2.16 % 2.16 %
test-other (2,939 utterances) 4.34 % 4.33 %

Long recordings are decoded in pause-aligned 25–35 s windows, the way Phonon-2's own engine does it; on concatenated LibriSpeech chapters that scores 1.79 % (test-clean) and 3.19 % (test-other). The engine's README has the protocol, the parity test against transformers and the speed table.

What is inside

Tensors Count Stored as
Encoder linears (feed-forward, fused Q/K/V, position, output, pointwise convs) 216 five-value: 2-bit signs + 1-bit lo/hi flags, [in, out], fp32 lo/hi per output, 229 MB
Subsampling linear, encoder projector 2 int8 with one fp32 scale per output column, exact from the container's int6
Embedding, LSTM, decoder projector, joint head 7 int8 with one fp32 scale per row, exact, 17 MB
Norms, biases, subsampling convs, depthwise conv with its BatchNorm folded in 354 fp32

Everything is exact except the BatchNorm fold, which is computed in double precision and stored as fp32.

Verifying and rebuilding

shasum -a 256 -c --ignore-missing SHA256SUMS.txt     # the archive, and the three files once unpacked

build.sh rebuilds the archive from Phonon-2's release and checks the result against those sums. It downloads phonon-2.bps.tar.zst (164 MB, sha256 98125795b6dda72f5c6eee9ba33d19815df65dcb18b50a357bf9f73c9935309e) from FermionResearch/Phonon-2, checks the container inside it (model.fermion, sha256 4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468, also recorded in weights.json under _meta), converts it with convert_weights.py and packs c_weights/ deterministically (ustar, fixed order, zero timestamps and owners, no gzip header name), so a rebuild gives this exact file:

python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt          # numpy, zstandard, huggingface_hub
./build.sh

convert_weights.py is a verbatim copy of the engine's converter, and fermion_container.py is Fermion Research's reference reader for the container, unchanged from the Phonon-2 repository.

Licence

The weights are CC-BY-4.0 (LICENSE-WEIGHTS-CC-BY-4.0.txt), the licence of Phonon-2 and of NVIDIA's Parakeet-TDT 0.6B v3 they derive from; the attribution chain, the changes made and the training-data terms are in NOTICE. convert_weights.py and build.sh are MIT (LICENSE-CODE-MIT.txt), as is the engine; fermion_container.py is Apache-2.0 (LICENSE-CODE-Apache-2.0.txt), Fermion Research.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eschmidbauer/phonon-2-c

Quantized
(7)
this model