Automatic Speech Recognition
GGUF
llama.cpp
asr
speech-recognition
streaming
real-time
low-latency
multilingual
quantized
conversational
Instructions to use FIT17/Confucius4-R2T2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FIT17/Confucius4-R2T2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FIT17/Confucius4-R2T2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/FIT17/Confucius4-R2T2-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use FIT17/Confucius4-R2T2-GGUF with Ollama:
ollama run hf.co/FIT17/Confucius4-R2T2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use FIT17/Confucius4-R2T2-GGUF with Docker Model Runner:
docker model run hf.co/FIT17/Confucius4-R2T2-GGUF:Q4_K_M
- Lemonade
How to use FIT17/Confucius4-R2T2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FIT17/Confucius4-R2T2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Confucius4-R2T2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
|
Download README.md from FIT17/Confucius4-R2T2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 6.26 kB
-
https://huggingface.co/FIT17/Confucius4-R2T2-GGUF/resolve/main/README.md
- Command line
-
hf download hf://FIT17/Confucius4-R2T2-GGUF/README.md
-
curl -L -o README.md https://huggingface.co/FIT17/Confucius4-R2T2-GGUF/resolve/main/README.md
6.26 kB
| license: other | |
| license_name: netease-model-use-license-agreement | |
| license_link: https://raw.githubusercontent.com/netease-youdao/Confucius4-R2T2/refs/heads/master/MODEL_LICENSE | |
| base_model: | |
| - netease-youdao/Confucius4-R2T2 | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - asr | |
| - speech-recognition | |
| - streaming | |
| - real-time | |
| - low-latency | |
| - multilingual | |
| - quantized | |
| pipeline_tag: automatic-speech-recognition | |
| # Confucius4-R2T2 Β· GGUF | |
| GGUF builds of [netease-youdao/Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2), | |
| a streaming ASR model trained on top of Qwen3-ASR-1.7B, so you can run it on | |
| [llama.cpp](https://github.com/ggml-org/llama.cpp) β CPU, CUDA and Metal all work. | |
| **δΈζηοΌ[README.zh-CN.md](README.zh-CN.md)** | |
| > The Q4_K_M build matches the official PyTorch bf16 reference, takes a third of the | |
| > space, and runs straight on llama.cpp. | |
| --- | |
| ## Files | |
| | File | Size | What it is | | |
| |---|---|---| | |
| | `Confucius4-R2T2-Q4_K_M.gguf` | 1.03 GiB | LLM Decoder | | |
| | `mmproj-Confucius4-R2T2-Q4_K_M.gguf` | 211 MiB | Audio Encoder | | |
| | `Confucius4-R2T2-BF16.gguf` | 3.21 GiB | LLM Decoder, BF16 (reference) | | |
| | `mmproj-Confucius4-R2T2-F16.gguf` | 612 MiB | Audio Encoder, F16 | | |
| --- | |
| ## Accuracy | |
| ### True streaming results | |
| R2T2 is a streaming model, and streaming is what it is actually for. On zh + en 569: | |
| | Implementation | Weights | Mode | CER | | |
| |---|---|---|---| | |
| | Official PyTorch | bf16 | offline | 2.03% | | |
| | Official PyTorch | bf16 | streaming | 2.65% | | |
| | [EdgeSpeak](https://edgespeak.com/) engine | Q4_K_M GGUF | offline | 2.04% | | |
| | [EdgeSpeak](https://edgespeak.com/) engine | Q4_K_M GGUF | streaming | **2.19%** | | |
| > **llama.cpp cannot do streaming today.** Both `llama-mtmd-cli` and `llama-server` take | |
| > the whole clip in one shot; there is no streaming ASR path. The two streaming rows above | |
| > use the same Q4_K_M pair, running on the EdgeSpeak on-device engine. | |
| ### Offline results | |
| > Running these GGUF files on llama.cpp means offline, non-streaming decoding. | |
| All public datasets, no training-set overlap. Greedy decoding throughout, no sampling. | |
| Lower CER is better. | |
| | Dataset | Clips | Official PyTorch bf16 (reference) | BF16 GGUF | Q4_K_M | Q4_K_M Β· [EdgeSpeak](https://edgespeak.com/) engine | | |
| |---|---|---|---|---|------| | |
| | `seedtts_testset` (Chinese) | 269 | 1.42% | 1.29% | 1.63% | 1.58% | | |
| | `seedtts_testset` (English) | 186 | 1.74% | 1.84% | 1.64% | 1.54% | | |
| | `voxpopuli-cleaned-aa` (English) | 106 | 3.28% | 3.31% | 3.24% | 3.18% | | |
| | `librispeech-long` (English) | 8 | 2.77% | 3.08% | 2.77% | 2.46% | | |
| | **zh + en total** | **569** | **2.03%** | **2.01%** | **2.12%** | **2.04%** | | |
| --- | |
| ## Quick start | |
| ### llama-mtmd-cli | |
| ```bash | |
| llama-mtmd-cli \ | |
| -m Confucius4-R2T2-Q4_K_M.gguf \ | |
| --mmproj mmproj-Confucius4-R2T2-Q4_K_M.gguf \ | |
| --audio your.wav \ | |
| -p " " --temp 0 -n 256 -ngl 99 \ | |
| --jinja # β REQUIRED. See below. | |
| ``` | |
| ### llama-server (OpenAI-compatible) | |
| ```bash | |
| llama-server \ | |
| -m Confucius4-R2T2-Q4_K_M.gguf \ | |
| --mmproj mmproj-Confucius4-R2T2-Q4_K_M.gguf \ | |
| -ngl 99 --temp 0 | |
| ``` | |
| ```bash | |
| curl http://localhost:8080/v1/audio/transcriptions \ | |
| -F file=@your.wav -F model=whisper-1 | |
| ``` | |
| --- | |
| ## Three things that will bite you | |
| ### 1. `llama-mtmd-cli` needs `--jinja`. Without it you silently get a different prompt. | |
| The prompt this model was trained on carries an **unconditional `system` block**, even | |
| when it is empty: | |
| ``` | |
| <|im_start|>system | |
| <|im_end|> | |
| <|im_start|>user | |
| <|audio_start|><|audio_pad|><|audio_end|><|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| The embedded Jinja template in these GGUF files produces exactly that. But | |
| `llama.cpp` has two rendering paths, and the legacy one **does not interpret Jinja at | |
| all** β it pattern-matches the template string (`<|im_start|>` β ChatML) and then runs a | |
| hardcoded C++ renderer that emits a `system` turn only if the caller supplied one. | |
| `llama-mtmd-cli` defaults to that legacy path; `--jinja` switches it to the real one. | |
| - `llama-mtmd-cli` β **defaults to legacy. Pass `--jinja`.** (Or set `LLAMA_ARG_JINJA=1`.) | |
| - `llama-server` β defaults to Jinja **on**. Don't pass `--no-jinja`. | |
| There is no metadata-only workaround: no builtin `llama.cpp` template family emits an | |
| unconditional `system` block, so editing the template string cannot fix the legacy path. | |
| The failure is silent β you just get slightly worse output. | |
| ### 2. `-p ""` is not "no prompt". | |
| An empty `-p` puts `llama-mtmd-cli` into interactive mode. Use `-p " "` for batch runs. | |
| (This leaves one stray space before `<|im_end|>`; there is no way to avoid it from the | |
| CLI, and its measured effect is negligible.) | |
| ### 3. Pad the tail with silence, or you lose the last few characters. | |
| R2T2 uses a **Longest Stable Prefix** paradigm: it emits a `|` to mark "this much is | |
| safe to commit, I need more audio before I continue". If your audio ends mid-sentence β | |
| which it does, because recordings stop the moment speech stops β the model never gets | |
| the evidence it needs to commit the final characters, and the transcript ends in a bare | |
| `|`. | |
| **Stripping the `|` does not recover the text.** The missing words were never generated. | |
| What fixes it is appending silence: | |
| ```bash | |
| ffmpeg -i in.wav -af apad=pad_dur=0.3 -ac 1 -ar 16000 -c:a pcm_s16le padded.wav | |
| ``` | |
| Measured on 22 clips that reproduce the truncation: | |
| | Tail silence | Still truncated | | |
| |---|---| | |
| | 0 s | 22 / 22 | | |
| | 0.1 s | 2 / 22 | | |
| | 0.15 β 0.25 s | 1 / 22 | | |
| | **0.3 s** | **0 / 22** | | |
| Before / after on one of them: | |
| ``` | |
| 0 s : I still say that this John Doe person is a| | |
| 0.3 s : I still say that this John Doe person is a myth. | |
| ``` | |
| The official streaming client pads 0.5 s before signalling end-of-stream. 0.3 s was the | |
| first fully clean point in our sweep; use 0.5 s if you want margin. | |
| --- | |
| ## License | |
| These are quantized derivatives. The original model's license applies unchanged: | |
| **NetEase Model Use License Agreement** β | |
| <https://raw.githubusercontent.com/netease-youdao/Confucius4-R2T2/refs/heads/master/MODEL_LICENSE> | |
| Base model: [Qwen/Qwen3-ASR-1.7B](https://huggingface.co/Qwen/Qwen3-ASR-1.7B). | |
| ## Credits | |
| Model by [NetEase Youdao](https://github.com/netease-youdao/Confucius4-R2T2). | |
| GGUF conversion, calibration and evaluation by us. | |