Text-to-Speech
Transformers
Safetensors
PyTorch
English
Chinese
breeze
text-generation
speech-generation
voice-clone
voice-design
voice-direction
cuda
Instructions to use BreezeBlue/Breeze-TTS-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BreezeBlue/Breeze-TTS-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="BreezeBlue/Breeze-TTS-2")# Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("BreezeBlue/Breeze-TTS-2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Normalize README formatting
Browse files
README.md
CHANGED
|
@@ -26,32 +26,24 @@ tags:
|
|
| 26 |
<a href="https://x.com/BreezeBlueX"><img src="https://img.shields.io/badge/X-Follow%20BreezeBlue-000000?logo=x&logoColor=white" alt="X"></a>
|
| 27 |
</div>
|
| 28 |
|
| 29 |
-
|
| 30 |
> [!IMPORTANT]
|
| 31 |
> Source code is licensed under Apache 2.0. Breeze TTS 2 model weights, derivative models, and self-hosted outputs are for research and non-commercial use only. See [License](#license-and-responsible-use).
|
| 32 |
|
| 33 |
-
|
| 34 |
## ๐ฐ News
|
| 35 |
|
| 36 |
-
|
| 37 |
- **[2026.08.25]** ๐ We open-source [Breeze TTS 2](https://huggingface.co/BreezeBlue/breeze-tts-2) model weights and the [PyTorch inference code](https://github.com/breezeblue-ai/breeze-tts).
|
| 38 |
- **[2026.08.07]** ๐ฅ We release the TTS benchmark suite for [voice design](https://github.com/breezeblue-ai/tts-voice-design-benchmark), [voice direction](https://github.com/breezeblue-ai/TTS-Voice-Direction-Benchmark), and [latency evaluation](https://github.com/breezeblue-ai/TTS-Latency-Benchmark).
|
| 39 |
|
| 40 |
-
|
| 41 |
## ๐ Introduction
|
| 42 |
|
| 43 |
-
|
| 44 |
Breeze TTS 2 is an open-weight text-to-speech model built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard, while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following capability supports reference-free voice design and reference-guided voice direction, while ultra-low-latency streaming enables responsive, expressive interaction.
|
| 45 |
|
| 46 |
-
|
| 47 |
<div align="center">
|
| 48 |
<img src="assets/tts-elo-leaderboard.svg" alt="Text-to-speech models ranked by Artificial Analysis Elo score" width="100%">
|
| 49 |
</div>
|
| 50 |
|
| 51 |
-
|
| 52 |
## โจ Highlights
|
| 53 |
|
| 54 |
-
|
| 55 |
- ๐๏ธ **Voice Clone** โ Uses reference audio with its exact transcript to preserve timbre, rhythm, emotion, and style.
|
| 56 |
- ๐จ **Voice Design** โ Creates a distinctive voice from a natural-language description, without reference audio.
|
| 57 |
- ๐๏ธ **Voice Direction** โ Clones a voice from reference audio while steering tone, emotion, pace, and delivery.
|
|
@@ -61,67 +53,50 @@ Breeze TTS 2 is an open-weight text-to-speech model built for real-time interact
|
|
| 61 |
- ๐พ **GPU-Efficient** โ Eager inference uses approximately 7.7 GiB of GPU memory; a 12 GB GPU is the minimum recommended configuration.
|
| 62 |
- ๐ **Bilingual Support** โ Generates natural English and Chinese speech with a single model.
|
| 63 |
|
| 64 |
-
|
| 65 |
## ๐ Quick Start
|
| 66 |
|
| 67 |
-
|
| 68 |
### Requirements
|
| 69 |
|
| 70 |
-
|
| 71 |
- Linux and Python 3.10 or newer
|
| 72 |
- A CUDA-capable NVIDIA GPU
|
| 73 |
- GPU memory: approximately 7.7 GiB for eager inference or 14.4 GiB with `--fast-all`; use a 12 GB GPU for eager or a 24 GB GPU for the fast path
|
| 74 |
- The Breeze TTS 2 checkpoint
|
| 75 |
|
| 76 |
-
|
| 77 |
### Installation
|
| 78 |
|
| 79 |
-
|
| 80 |
Download the inference code:
|
| 81 |
|
| 82 |
-
|
| 83 |
```bash
|
| 84 |
git clone https://github.com/breezeblue-ai/breeze-tts.git
|
| 85 |
cd breeze-tts
|
| 86 |
```
|
| 87 |
|
| 88 |
-
|
| 89 |
Install the dependencies:
|
| 90 |
|
| 91 |
-
|
| 92 |
```bash
|
| 93 |
python -m pip install -r requirements.txt
|
| 94 |
```
|
| 95 |
|
| 96 |
-
|
| 97 |
All required model components are included in the Breeze TTS 2 checkpoint.
|
| 98 |
|
| 99 |
-
|
| 100 |
For the tested CUDA environment, build the included Docker image:
|
| 101 |
|
| 102 |
-
|
| 103 |
```bash
|
| 104 |
bash docker/build.sh
|
| 105 |
```
|
| 106 |
|
| 107 |
-
|
| 108 |
The default image targets H100/Hopper (sm90). For A100:
|
| 109 |
|
| 110 |
-
|
| 111 |
```bash
|
| 112 |
FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh
|
| 113 |
```
|
| 114 |
|
| 115 |
-
|
| 116 |
### ๐๏ธ Voice Clone
|
| 117 |
|
| 118 |
-
|
| 119 |
Clone a speaker from clean reference audio and its exact transcript.
|
| 120 |
|
| 121 |
-
|
| 122 |
#### English
|
| 123 |
|
| 124 |
-
|
| 125 |
```bash
|
| 126 |
python infer.py ../breeze-tts-2 \
|
| 127 |
--ref-audio reference_en.wav \
|
|
@@ -130,10 +105,8 @@ python infer.py ../breeze-tts-2 \
|
|
| 130 |
--output outputs/voice_clone_en.wav
|
| 131 |
```
|
| 132 |
|
| 133 |
-
|
| 134 |
#### Chinese
|
| 135 |
|
| 136 |
-
|
| 137 |
```bash
|
| 138 |
python infer.py ../breeze-tts-2 \
|
| 139 |
--ref-audio reference_zh.wav \
|
|
@@ -142,19 +115,14 @@ python infer.py ../breeze-tts-2 \
|
|
| 142 |
--output outputs/voice_clone_zh.wav
|
| 143 |
```
|
| 144 |
|
| 145 |
-
|
| 146 |
Reference audio should contain clean speech with minimal background noise.
|
| 147 |
|
| 148 |
-
|
| 149 |
### ๐จ Voice Design
|
| 150 |
|
| 151 |
-
|
| 152 |
Create a voice from a natural-language description without reference audio. Match the instruction language to the target text. Use `--cfg-scale 4` to strengthen instruction-following.
|
| 153 |
|
| 154 |
-
|
| 155 |
#### English
|
| 156 |
|
| 157 |
-
|
| 158 |
```bash
|
| 159 |
python infer.py ../breeze-tts-2 \
|
| 160 |
--text "(sigh) Welcome aboard. Your journey begins now." \
|
|
@@ -163,10 +131,8 @@ python infer.py ../breeze-tts-2 \
|
|
| 163 |
--output outputs/voice_design_en.wav
|
| 164 |
```
|
| 165 |
|
| 166 |
-
|
| 167 |
#### Chinese
|
| 168 |
|
| 169 |
-
|
| 170 |
```bash
|
| 171 |
python infer.py ../breeze-tts-2 \
|
| 172 |
--text "[็ฌ] ๆฌข่ฟๆฅๅฐไปๆ็ๆ
ไบๆถ้ด๏ผ่ฎฉๆไปฌไธ่ตทๅผๅงๅงใ" \
|
|
@@ -175,13 +141,10 @@ python infer.py ../breeze-tts-2 \
|
|
| 175 |
--output outputs/voice_design_zh.wav
|
| 176 |
```
|
| 177 |
|
| 178 |
-
|
| 179 |
### ๐๏ธ Voice Direction
|
| 180 |
|
| 181 |
-
|
| 182 |
Keep the identity of a reference speaker while directing tone, emotion, pace, and delivery. Use `--cfg-scale 4` to strengthen instruction-following.
|
| 183 |
|
| 184 |
-
|
| 185 |
```bash
|
| 186 |
python infer.py ../breeze-tts-2 \
|
| 187 |
--ref-audio reference.wav \
|
|
@@ -192,21 +155,16 @@ python infer.py ../breeze-tts-2 \
|
|
| 192 |
--output outputs/voice_direction.wav
|
| 193 |
```
|
| 194 |
|
| 195 |
-
|
| 196 |
### ๐ Streaming API
|
| 197 |
|
| 198 |
-
|
| 199 |
Start the single-concurrency streaming API. It uses the same PyTorch runtime and eager execution by default:
|
| 200 |
|
| 201 |
-
|
| 202 |
```bash
|
| 203 |
python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860
|
| 204 |
```
|
| 205 |
|
| 206 |
-
|
| 207 |
Send a Voice Direction request with reference audio and CFG 4:
|
| 208 |
|
| 209 |
-
|
| 210 |
```bash
|
| 211 |
curl -X POST http://127.0.0.1:7860/v1/audio/speech \
|
| 212 |
-F "cfg_scale=4" \
|
|
@@ -218,16 +176,12 @@ curl -X POST http://127.0.0.1:7860/v1/audio/speech \
|
|
| 218 |
--output voice_direction.pcm
|
| 219 |
```
|
| 220 |
|
| 221 |
-
|
| 222 |
The response is streaming mono 24 kHz signed 16-bit little-endian PCM. Start the API with `--fast-all` to enable the fast path.
|
| 223 |
|
| 224 |
-
|
| 225 |
### โก Fast Inference Options
|
| 226 |
|
| 227 |
-
|
| 228 |
Both the CLI and API use eager streaming by default and skip graph warmup. Pass `--fast-all` to enable the best configuration for every inference stage when the additional cold-start time is acceptable. Each stage can also be controlled independently:
|
| 229 |
|
| 230 |
-
|
| 231 |
| Stage | Fast parameter | Disabled | Enabled |
|
| 232 |
| --- | --- | --- | --- |
|
| 233 |
| Text encoder | `--[no-]fast-text-encoder` | Native eager forward | Static CUDA Graph selected by CFG shape and text-length bucket |
|
|
@@ -236,23 +190,15 @@ Both the CLI and API use eager streaming by default and skip graph warmup. Pass
|
|
| 236 |
| Depth decoder | `--[no-]fast-depth-decoder` | Native eager depth loop | Full-graph compilation with CFG-shape CUDA Graphs |
|
| 237 |
| Codec | `--[no-]fast-codec` | Eager streaming decode | Single-request streaming CUDA Graph with one-frame chunks |
|
| 238 |
|
| 239 |
-
|
| 240 |
Individual stage flags are intended for profiling and debugging.
|
| 241 |
|
| 242 |
|
| 243 |
-
|
| 244 |
-
|
| 245 |
## License and Responsible Use
|
| 246 |
|
| 247 |
-
|
| 248 |
The source code is licensed under the [Apache License, Version 2.0](https://github.com/breezeblue-ai/breeze-tts/blob/main/LICENSE). The audio tokenizer is based on [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) by the Alibaba Qwen Team and is licensed under the Apache License, Version 2.0. Model weights, checkpoints, adapters, derivative models, and self-hosted outputs are governed separately by the [BreezeBlue Research and Non-Commercial License](https://huggingface.co/BreezeBlue/Breeze-TTS-2/blob/main/LICENSE). The Apache License does not grant rights to use the model commercially.
|
| 249 |
|
| 250 |
-
|
| 251 |
If you have an active paid subscription, outputs you generate through BreezeBlue's hosted platform or API at [breezeblue.ai](https://breezeblue.ai/) can be used commercially, subject to our [Terms of Service](https://breezeblue.ai/legal/terms). A paid subscription does not grant commercial rights to the open-weight model or self-hosted outputs.
|
| 252 |
|
| 253 |
-
|
| 254 |
You are responsible for complying with applicable laws and obtaining all necessary rights and consents for inputs, reference audio, voices, and outputs. Unauthorized voice cloning, impersonation, fraud, and other unlawful or harmful uses are prohibited.
|
| 255 |
|
| 256 |
-
|
| 257 |
The code and Model Materials are provided "AS IS," without warranties or liability to the maximum extent permitted by law. Third-party components remain subject to their respective licenses.
|
| 258 |
-
|
|
|
|
| 26 |
<a href="https://x.com/BreezeBlueX"><img src="https://img.shields.io/badge/X-Follow%20BreezeBlue-000000?logo=x&logoColor=white" alt="X"></a>
|
| 27 |
</div>
|
| 28 |
|
|
|
|
| 29 |
> [!IMPORTANT]
|
| 30 |
> Source code is licensed under Apache 2.0. Breeze TTS 2 model weights, derivative models, and self-hosted outputs are for research and non-commercial use only. See [License](#license-and-responsible-use).
|
| 31 |
|
|
|
|
| 32 |
## ๐ฐ News
|
| 33 |
|
|
|
|
| 34 |
- **[2026.08.25]** ๐ We open-source [Breeze TTS 2](https://huggingface.co/BreezeBlue/breeze-tts-2) model weights and the [PyTorch inference code](https://github.com/breezeblue-ai/breeze-tts).
|
| 35 |
- **[2026.08.07]** ๐ฅ We release the TTS benchmark suite for [voice design](https://github.com/breezeblue-ai/tts-voice-design-benchmark), [voice direction](https://github.com/breezeblue-ai/TTS-Voice-Direction-Benchmark), and [latency evaluation](https://github.com/breezeblue-ai/TTS-Latency-Benchmark).
|
| 36 |
|
|
|
|
| 37 |
## ๐ Introduction
|
| 38 |
|
|
|
|
| 39 |
Breeze TTS 2 is an open-weight text-to-speech model built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard, while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following capability supports reference-free voice design and reference-guided voice direction, while ultra-low-latency streaming enables responsive, expressive interaction.
|
| 40 |
|
|
|
|
| 41 |
<div align="center">
|
| 42 |
<img src="assets/tts-elo-leaderboard.svg" alt="Text-to-speech models ranked by Artificial Analysis Elo score" width="100%">
|
| 43 |
</div>
|
| 44 |
|
|
|
|
| 45 |
## โจ Highlights
|
| 46 |
|
|
|
|
| 47 |
- ๐๏ธ **Voice Clone** โ Uses reference audio with its exact transcript to preserve timbre, rhythm, emotion, and style.
|
| 48 |
- ๐จ **Voice Design** โ Creates a distinctive voice from a natural-language description, without reference audio.
|
| 49 |
- ๐๏ธ **Voice Direction** โ Clones a voice from reference audio while steering tone, emotion, pace, and delivery.
|
|
|
|
| 53 |
- ๐พ **GPU-Efficient** โ Eager inference uses approximately 7.7 GiB of GPU memory; a 12 GB GPU is the minimum recommended configuration.
|
| 54 |
- ๐ **Bilingual Support** โ Generates natural English and Chinese speech with a single model.
|
| 55 |
|
|
|
|
| 56 |
## ๐ Quick Start
|
| 57 |
|
|
|
|
| 58 |
### Requirements
|
| 59 |
|
|
|
|
| 60 |
- Linux and Python 3.10 or newer
|
| 61 |
- A CUDA-capable NVIDIA GPU
|
| 62 |
- GPU memory: approximately 7.7 GiB for eager inference or 14.4 GiB with `--fast-all`; use a 12 GB GPU for eager or a 24 GB GPU for the fast path
|
| 63 |
- The Breeze TTS 2 checkpoint
|
| 64 |
|
|
|
|
| 65 |
### Installation
|
| 66 |
|
|
|
|
| 67 |
Download the inference code:
|
| 68 |
|
|
|
|
| 69 |
```bash
|
| 70 |
git clone https://github.com/breezeblue-ai/breeze-tts.git
|
| 71 |
cd breeze-tts
|
| 72 |
```
|
| 73 |
|
|
|
|
| 74 |
Install the dependencies:
|
| 75 |
|
|
|
|
| 76 |
```bash
|
| 77 |
python -m pip install -r requirements.txt
|
| 78 |
```
|
| 79 |
|
|
|
|
| 80 |
All required model components are included in the Breeze TTS 2 checkpoint.
|
| 81 |
|
|
|
|
| 82 |
For the tested CUDA environment, build the included Docker image:
|
| 83 |
|
|
|
|
| 84 |
```bash
|
| 85 |
bash docker/build.sh
|
| 86 |
```
|
| 87 |
|
|
|
|
| 88 |
The default image targets H100/Hopper (sm90). For A100:
|
| 89 |
|
|
|
|
| 90 |
```bash
|
| 91 |
FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh
|
| 92 |
```
|
| 93 |
|
|
|
|
| 94 |
### ๐๏ธ Voice Clone
|
| 95 |
|
|
|
|
| 96 |
Clone a speaker from clean reference audio and its exact transcript.
|
| 97 |
|
|
|
|
| 98 |
#### English
|
| 99 |
|
|
|
|
| 100 |
```bash
|
| 101 |
python infer.py ../breeze-tts-2 \
|
| 102 |
--ref-audio reference_en.wav \
|
|
|
|
| 105 |
--output outputs/voice_clone_en.wav
|
| 106 |
```
|
| 107 |
|
|
|
|
| 108 |
#### Chinese
|
| 109 |
|
|
|
|
| 110 |
```bash
|
| 111 |
python infer.py ../breeze-tts-2 \
|
| 112 |
--ref-audio reference_zh.wav \
|
|
|
|
| 115 |
--output outputs/voice_clone_zh.wav
|
| 116 |
```
|
| 117 |
|
|
|
|
| 118 |
Reference audio should contain clean speech with minimal background noise.
|
| 119 |
|
|
|
|
| 120 |
### ๐จ Voice Design
|
| 121 |
|
|
|
|
| 122 |
Create a voice from a natural-language description without reference audio. Match the instruction language to the target text. Use `--cfg-scale 4` to strengthen instruction-following.
|
| 123 |
|
|
|
|
| 124 |
#### English
|
| 125 |
|
|
|
|
| 126 |
```bash
|
| 127 |
python infer.py ../breeze-tts-2 \
|
| 128 |
--text "(sigh) Welcome aboard. Your journey begins now." \
|
|
|
|
| 131 |
--output outputs/voice_design_en.wav
|
| 132 |
```
|
| 133 |
|
|
|
|
| 134 |
#### Chinese
|
| 135 |
|
|
|
|
| 136 |
```bash
|
| 137 |
python infer.py ../breeze-tts-2 \
|
| 138 |
--text "[็ฌ] ๆฌข่ฟๆฅๅฐไปๆ็ๆ
ไบๆถ้ด๏ผ่ฎฉๆไปฌไธ่ตทๅผๅงๅงใ" \
|
|
|
|
| 141 |
--output outputs/voice_design_zh.wav
|
| 142 |
```
|
| 143 |
|
|
|
|
| 144 |
### ๐๏ธ Voice Direction
|
| 145 |
|
|
|
|
| 146 |
Keep the identity of a reference speaker while directing tone, emotion, pace, and delivery. Use `--cfg-scale 4` to strengthen instruction-following.
|
| 147 |
|
|
|
|
| 148 |
```bash
|
| 149 |
python infer.py ../breeze-tts-2 \
|
| 150 |
--ref-audio reference.wav \
|
|
|
|
| 155 |
--output outputs/voice_direction.wav
|
| 156 |
```
|
| 157 |
|
|
|
|
| 158 |
### ๐ Streaming API
|
| 159 |
|
|
|
|
| 160 |
Start the single-concurrency streaming API. It uses the same PyTorch runtime and eager execution by default:
|
| 161 |
|
|
|
|
| 162 |
```bash
|
| 163 |
python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860
|
| 164 |
```
|
| 165 |
|
|
|
|
| 166 |
Send a Voice Direction request with reference audio and CFG 4:
|
| 167 |
|
|
|
|
| 168 |
```bash
|
| 169 |
curl -X POST http://127.0.0.1:7860/v1/audio/speech \
|
| 170 |
-F "cfg_scale=4" \
|
|
|
|
| 176 |
--output voice_direction.pcm
|
| 177 |
```
|
| 178 |
|
|
|
|
| 179 |
The response is streaming mono 24 kHz signed 16-bit little-endian PCM. Start the API with `--fast-all` to enable the fast path.
|
| 180 |
|
|
|
|
| 181 |
### โก Fast Inference Options
|
| 182 |
|
|
|
|
| 183 |
Both the CLI and API use eager streaming by default and skip graph warmup. Pass `--fast-all` to enable the best configuration for every inference stage when the additional cold-start time is acceptable. Each stage can also be controlled independently:
|
| 184 |
|
|
|
|
| 185 |
| Stage | Fast parameter | Disabled | Enabled |
|
| 186 |
| --- | --- | --- | --- |
|
| 187 |
| Text encoder | `--[no-]fast-text-encoder` | Native eager forward | Static CUDA Graph selected by CFG shape and text-length bucket |
|
|
|
|
| 190 |
| Depth decoder | `--[no-]fast-depth-decoder` | Native eager depth loop | Full-graph compilation with CFG-shape CUDA Graphs |
|
| 191 |
| Codec | `--[no-]fast-codec` | Eager streaming decode | Single-request streaming CUDA Graph with one-frame chunks |
|
| 192 |
|
|
|
|
| 193 |
Individual stage flags are intended for profiling and debugging.
|
| 194 |
|
| 195 |
|
|
|
|
|
|
|
| 196 |
## License and Responsible Use
|
| 197 |
|
|
|
|
| 198 |
The source code is licensed under the [Apache License, Version 2.0](https://github.com/breezeblue-ai/breeze-tts/blob/main/LICENSE). The audio tokenizer is based on [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) by the Alibaba Qwen Team and is licensed under the Apache License, Version 2.0. Model weights, checkpoints, adapters, derivative models, and self-hosted outputs are governed separately by the [BreezeBlue Research and Non-Commercial License](https://huggingface.co/BreezeBlue/Breeze-TTS-2/blob/main/LICENSE). The Apache License does not grant rights to use the model commercially.
|
| 199 |
|
|
|
|
| 200 |
If you have an active paid subscription, outputs you generate through BreezeBlue's hosted platform or API at [breezeblue.ai](https://breezeblue.ai/) can be used commercially, subject to our [Terms of Service](https://breezeblue.ai/legal/terms). A paid subscription does not grant commercial rights to the open-weight model or self-hosted outputs.
|
| 201 |
|
|
|
|
| 202 |
You are responsible for complying with applicable laws and obtaining all necessary rights and consents for inputs, reference audio, voices, and outputs. Unauthorized voice cloning, impersonation, fraud, and other unlawful or harmful uses are prohibited.
|
| 203 |
|
|
|
|
| 204 |
The code and Model Materials are provided "AS IS," without warranties or liability to the maximum extent permitted by law. Third-party components remain subject to their respective licenses.
|
|
|