byangderek commited on
Commit
799624c
ยท
verified ยท
1 Parent(s): ff883d0

Normalize README formatting

Browse files
Files changed (1) hide show
  1. README.md +0 -54
README.md CHANGED
@@ -26,32 +26,24 @@ tags:
26
  <a href="https://x.com/BreezeBlueX"><img src="https://img.shields.io/badge/X-Follow%20BreezeBlue-000000?logo=x&logoColor=white" alt="X"></a>
27
  </div>
28
 
29
-
30
  > [!IMPORTANT]
31
  > Source code is licensed under Apache 2.0. Breeze TTS 2 model weights, derivative models, and self-hosted outputs are for research and non-commercial use only. See [License](#license-and-responsible-use).
32
 
33
-
34
  ## ๐Ÿ“ฐ News
35
 
36
-
37
  - **[2026.08.25]** ๐ŸŽ‰ We open-source [Breeze TTS 2](https://huggingface.co/BreezeBlue/breeze-tts-2) model weights and the [PyTorch inference code](https://github.com/breezeblue-ai/breeze-tts).
38
  - **[2026.08.07]** ๐Ÿ”ฅ We release the TTS benchmark suite for [voice design](https://github.com/breezeblue-ai/tts-voice-design-benchmark), [voice direction](https://github.com/breezeblue-ai/TTS-Voice-Direction-Benchmark), and [latency evaluation](https://github.com/breezeblue-ai/TTS-Latency-Benchmark).
39
 
40
-
41
  ## ๐Ÿ“– Introduction
42
 
43
-
44
  Breeze TTS 2 is an open-weight text-to-speech model built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard, while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following capability supports reference-free voice design and reference-guided voice direction, while ultra-low-latency streaming enables responsive, expressive interaction.
45
 
46
-
47
  <div align="center">
48
  <img src="assets/tts-elo-leaderboard.svg" alt="Text-to-speech models ranked by Artificial Analysis Elo score" width="100%">
49
  </div>
50
 
51
-
52
  ## โœจ Highlights
53
 
54
-
55
  - ๐ŸŽ™๏ธ **Voice Clone** โ€” Uses reference audio with its exact transcript to preserve timbre, rhythm, emotion, and style.
56
  - ๐ŸŽจ **Voice Design** โ€” Creates a distinctive voice from a natural-language description, without reference audio.
57
  - ๐ŸŽ›๏ธ **Voice Direction** โ€” Clones a voice from reference audio while steering tone, emotion, pace, and delivery.
@@ -61,67 +53,50 @@ Breeze TTS 2 is an open-weight text-to-speech model built for real-time interact
61
  - ๐Ÿ’พ **GPU-Efficient** โ€” Eager inference uses approximately 7.7 GiB of GPU memory; a 12 GB GPU is the minimum recommended configuration.
62
  - ๐ŸŒ **Bilingual Support** โ€” Generates natural English and Chinese speech with a single model.
63
 
64
-
65
  ## ๐Ÿš€ Quick Start
66
 
67
-
68
  ### Requirements
69
 
70
-
71
  - Linux and Python 3.10 or newer
72
  - A CUDA-capable NVIDIA GPU
73
  - GPU memory: approximately 7.7 GiB for eager inference or 14.4 GiB with `--fast-all`; use a 12 GB GPU for eager or a 24 GB GPU for the fast path
74
  - The Breeze TTS 2 checkpoint
75
 
76
-
77
  ### Installation
78
 
79
-
80
  Download the inference code:
81
 
82
-
83
  ```bash
84
  git clone https://github.com/breezeblue-ai/breeze-tts.git
85
  cd breeze-tts
86
  ```
87
 
88
-
89
  Install the dependencies:
90
 
91
-
92
  ```bash
93
  python -m pip install -r requirements.txt
94
  ```
95
 
96
-
97
  All required model components are included in the Breeze TTS 2 checkpoint.
98
 
99
-
100
  For the tested CUDA environment, build the included Docker image:
101
 
102
-
103
  ```bash
104
  bash docker/build.sh
105
  ```
106
 
107
-
108
  The default image targets H100/Hopper (sm90). For A100:
109
 
110
-
111
  ```bash
112
  FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh
113
  ```
114
 
115
-
116
  ### ๐ŸŽ™๏ธ Voice Clone
117
 
118
-
119
  Clone a speaker from clean reference audio and its exact transcript.
120
 
121
-
122
  #### English
123
 
124
-
125
  ```bash
126
  python infer.py ../breeze-tts-2 \
127
  --ref-audio reference_en.wav \
@@ -130,10 +105,8 @@ python infer.py ../breeze-tts-2 \
130
  --output outputs/voice_clone_en.wav
131
  ```
132
 
133
-
134
  #### Chinese
135
 
136
-
137
  ```bash
138
  python infer.py ../breeze-tts-2 \
139
  --ref-audio reference_zh.wav \
@@ -142,19 +115,14 @@ python infer.py ../breeze-tts-2 \
142
  --output outputs/voice_clone_zh.wav
143
  ```
144
 
145
-
146
  Reference audio should contain clean speech with minimal background noise.
147
 
148
-
149
  ### ๐ŸŽจ Voice Design
150
 
151
-
152
  Create a voice from a natural-language description without reference audio. Match the instruction language to the target text. Use `--cfg-scale 4` to strengthen instruction-following.
153
 
154
-
155
  #### English
156
 
157
-
158
  ```bash
159
  python infer.py ../breeze-tts-2 \
160
  --text "(sigh) Welcome aboard. Your journey begins now." \
@@ -163,10 +131,8 @@ python infer.py ../breeze-tts-2 \
163
  --output outputs/voice_design_en.wav
164
  ```
165
 
166
-
167
  #### Chinese
168
 
169
-
170
  ```bash
171
  python infer.py ../breeze-tts-2 \
172
  --text "[็ฌ‘] ๆฌข่ฟŽๆฅๅˆฐไปŠๆ™š็š„ๆ•…ไบ‹ๆ—ถ้—ด๏ผŒ่ฎฉๆˆ‘ไปฌไธ€่ตทๅผ€ๅง‹ๅงใ€‚" \
@@ -175,13 +141,10 @@ python infer.py ../breeze-tts-2 \
175
  --output outputs/voice_design_zh.wav
176
  ```
177
 
178
-
179
  ### ๐ŸŽ›๏ธ Voice Direction
180
 
181
-
182
  Keep the identity of a reference speaker while directing tone, emotion, pace, and delivery. Use `--cfg-scale 4` to strengthen instruction-following.
183
 
184
-
185
  ```bash
186
  python infer.py ../breeze-tts-2 \
187
  --ref-audio reference.wav \
@@ -192,21 +155,16 @@ python infer.py ../breeze-tts-2 \
192
  --output outputs/voice_direction.wav
193
  ```
194
 
195
-
196
  ### ๐ŸŒ Streaming API
197
 
198
-
199
  Start the single-concurrency streaming API. It uses the same PyTorch runtime and eager execution by default:
200
 
201
-
202
  ```bash
203
  python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860
204
  ```
205
 
206
-
207
  Send a Voice Direction request with reference audio and CFG 4:
208
 
209
-
210
  ```bash
211
  curl -X POST http://127.0.0.1:7860/v1/audio/speech \
212
  -F "cfg_scale=4" \
@@ -218,16 +176,12 @@ curl -X POST http://127.0.0.1:7860/v1/audio/speech \
218
  --output voice_direction.pcm
219
  ```
220
 
221
-
222
  The response is streaming mono 24 kHz signed 16-bit little-endian PCM. Start the API with `--fast-all` to enable the fast path.
223
 
224
-
225
  ### โšก Fast Inference Options
226
 
227
-
228
  Both the CLI and API use eager streaming by default and skip graph warmup. Pass `--fast-all` to enable the best configuration for every inference stage when the additional cold-start time is acceptable. Each stage can also be controlled independently:
229
 
230
-
231
  | Stage | Fast parameter | Disabled | Enabled |
232
  | --- | --- | --- | --- |
233
  | Text encoder | `--[no-]fast-text-encoder` | Native eager forward | Static CUDA Graph selected by CFG shape and text-length bucket |
@@ -236,23 +190,15 @@ Both the CLI and API use eager streaming by default and skip graph warmup. Pass
236
  | Depth decoder | `--[no-]fast-depth-decoder` | Native eager depth loop | Full-graph compilation with CFG-shape CUDA Graphs |
237
  | Codec | `--[no-]fast-codec` | Eager streaming decode | Single-request streaming CUDA Graph with one-frame chunks |
238
 
239
-
240
  Individual stage flags are intended for profiling and debugging.
241
 
242
 
243
-
244
-
245
  ## License and Responsible Use
246
 
247
-
248
  The source code is licensed under the [Apache License, Version 2.0](https://github.com/breezeblue-ai/breeze-tts/blob/main/LICENSE). The audio tokenizer is based on [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) by the Alibaba Qwen Team and is licensed under the Apache License, Version 2.0. Model weights, checkpoints, adapters, derivative models, and self-hosted outputs are governed separately by the [BreezeBlue Research and Non-Commercial License](https://huggingface.co/BreezeBlue/Breeze-TTS-2/blob/main/LICENSE). The Apache License does not grant rights to use the model commercially.
249
 
250
-
251
  If you have an active paid subscription, outputs you generate through BreezeBlue's hosted platform or API at [breezeblue.ai](https://breezeblue.ai/) can be used commercially, subject to our [Terms of Service](https://breezeblue.ai/legal/terms). A paid subscription does not grant commercial rights to the open-weight model or self-hosted outputs.
252
 
253
-
254
  You are responsible for complying with applicable laws and obtaining all necessary rights and consents for inputs, reference audio, voices, and outputs. Unauthorized voice cloning, impersonation, fraud, and other unlawful or harmful uses are prohibited.
255
 
256
-
257
  The code and Model Materials are provided "AS IS," without warranties or liability to the maximum extent permitted by law. Third-party components remain subject to their respective licenses.
258
-
 
26
  <a href="https://x.com/BreezeBlueX"><img src="https://img.shields.io/badge/X-Follow%20BreezeBlue-000000?logo=x&logoColor=white" alt="X"></a>
27
  </div>
28
 
 
29
  > [!IMPORTANT]
30
  > Source code is licensed under Apache 2.0. Breeze TTS 2 model weights, derivative models, and self-hosted outputs are for research and non-commercial use only. See [License](#license-and-responsible-use).
31
 
 
32
  ## ๐Ÿ“ฐ News
33
 
 
34
  - **[2026.08.25]** ๐ŸŽ‰ We open-source [Breeze TTS 2](https://huggingface.co/BreezeBlue/breeze-tts-2) model weights and the [PyTorch inference code](https://github.com/breezeblue-ai/breeze-tts).
35
  - **[2026.08.07]** ๐Ÿ”ฅ We release the TTS benchmark suite for [voice design](https://github.com/breezeblue-ai/tts-voice-design-benchmark), [voice direction](https://github.com/breezeblue-ai/TTS-Voice-Direction-Benchmark), and [latency evaluation](https://github.com/breezeblue-ai/TTS-Latency-Benchmark).
36
 
 
37
  ## ๐Ÿ“– Introduction
38
 
 
39
  Breeze TTS 2 is an open-weight text-to-speech model built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard, while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following capability supports reference-free voice design and reference-guided voice direction, while ultra-low-latency streaming enables responsive, expressive interaction.
40
 
 
41
  <div align="center">
42
  <img src="assets/tts-elo-leaderboard.svg" alt="Text-to-speech models ranked by Artificial Analysis Elo score" width="100%">
43
  </div>
44
 
 
45
  ## โœจ Highlights
46
 
 
47
  - ๐ŸŽ™๏ธ **Voice Clone** โ€” Uses reference audio with its exact transcript to preserve timbre, rhythm, emotion, and style.
48
  - ๐ŸŽจ **Voice Design** โ€” Creates a distinctive voice from a natural-language description, without reference audio.
49
  - ๐ŸŽ›๏ธ **Voice Direction** โ€” Clones a voice from reference audio while steering tone, emotion, pace, and delivery.
 
53
  - ๐Ÿ’พ **GPU-Efficient** โ€” Eager inference uses approximately 7.7 GiB of GPU memory; a 12 GB GPU is the minimum recommended configuration.
54
  - ๐ŸŒ **Bilingual Support** โ€” Generates natural English and Chinese speech with a single model.
55
 
 
56
  ## ๐Ÿš€ Quick Start
57
 
 
58
  ### Requirements
59
 
 
60
  - Linux and Python 3.10 or newer
61
  - A CUDA-capable NVIDIA GPU
62
  - GPU memory: approximately 7.7 GiB for eager inference or 14.4 GiB with `--fast-all`; use a 12 GB GPU for eager or a 24 GB GPU for the fast path
63
  - The Breeze TTS 2 checkpoint
64
 
 
65
  ### Installation
66
 
 
67
  Download the inference code:
68
 
 
69
  ```bash
70
  git clone https://github.com/breezeblue-ai/breeze-tts.git
71
  cd breeze-tts
72
  ```
73
 
 
74
  Install the dependencies:
75
 
 
76
  ```bash
77
  python -m pip install -r requirements.txt
78
  ```
79
 
 
80
  All required model components are included in the Breeze TTS 2 checkpoint.
81
 
 
82
  For the tested CUDA environment, build the included Docker image:
83
 
 
84
  ```bash
85
  bash docker/build.sh
86
  ```
87
 
 
88
  The default image targets H100/Hopper (sm90). For A100:
89
 
 
90
  ```bash
91
  FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh
92
  ```
93
 
 
94
  ### ๐ŸŽ™๏ธ Voice Clone
95
 
 
96
  Clone a speaker from clean reference audio and its exact transcript.
97
 
 
98
  #### English
99
 
 
100
  ```bash
101
  python infer.py ../breeze-tts-2 \
102
  --ref-audio reference_en.wav \
 
105
  --output outputs/voice_clone_en.wav
106
  ```
107
 
 
108
  #### Chinese
109
 
 
110
  ```bash
111
  python infer.py ../breeze-tts-2 \
112
  --ref-audio reference_zh.wav \
 
115
  --output outputs/voice_clone_zh.wav
116
  ```
117
 
 
118
  Reference audio should contain clean speech with minimal background noise.
119
 
 
120
  ### ๐ŸŽจ Voice Design
121
 
 
122
  Create a voice from a natural-language description without reference audio. Match the instruction language to the target text. Use `--cfg-scale 4` to strengthen instruction-following.
123
 
 
124
  #### English
125
 
 
126
  ```bash
127
  python infer.py ../breeze-tts-2 \
128
  --text "(sigh) Welcome aboard. Your journey begins now." \
 
131
  --output outputs/voice_design_en.wav
132
  ```
133
 
 
134
  #### Chinese
135
 
 
136
  ```bash
137
  python infer.py ../breeze-tts-2 \
138
  --text "[็ฌ‘] ๆฌข่ฟŽๆฅๅˆฐไปŠๆ™š็š„ๆ•…ไบ‹ๆ—ถ้—ด๏ผŒ่ฎฉๆˆ‘ไปฌไธ€่ตทๅผ€ๅง‹ๅงใ€‚" \
 
141
  --output outputs/voice_design_zh.wav
142
  ```
143
 
 
144
  ### ๐ŸŽ›๏ธ Voice Direction
145
 
 
146
  Keep the identity of a reference speaker while directing tone, emotion, pace, and delivery. Use `--cfg-scale 4` to strengthen instruction-following.
147
 
 
148
  ```bash
149
  python infer.py ../breeze-tts-2 \
150
  --ref-audio reference.wav \
 
155
  --output outputs/voice_direction.wav
156
  ```
157
 
 
158
  ### ๐ŸŒ Streaming API
159
 
 
160
  Start the single-concurrency streaming API. It uses the same PyTorch runtime and eager execution by default:
161
 
 
162
  ```bash
163
  python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860
164
  ```
165
 
 
166
  Send a Voice Direction request with reference audio and CFG 4:
167
 
 
168
  ```bash
169
  curl -X POST http://127.0.0.1:7860/v1/audio/speech \
170
  -F "cfg_scale=4" \
 
176
  --output voice_direction.pcm
177
  ```
178
 
 
179
  The response is streaming mono 24 kHz signed 16-bit little-endian PCM. Start the API with `--fast-all` to enable the fast path.
180
 
 
181
  ### โšก Fast Inference Options
182
 
 
183
  Both the CLI and API use eager streaming by default and skip graph warmup. Pass `--fast-all` to enable the best configuration for every inference stage when the additional cold-start time is acceptable. Each stage can also be controlled independently:
184
 
 
185
  | Stage | Fast parameter | Disabled | Enabled |
186
  | --- | --- | --- | --- |
187
  | Text encoder | `--[no-]fast-text-encoder` | Native eager forward | Static CUDA Graph selected by CFG shape and text-length bucket |
 
190
  | Depth decoder | `--[no-]fast-depth-decoder` | Native eager depth loop | Full-graph compilation with CFG-shape CUDA Graphs |
191
  | Codec | `--[no-]fast-codec` | Eager streaming decode | Single-request streaming CUDA Graph with one-frame chunks |
192
 
 
193
  Individual stage flags are intended for profiling and debugging.
194
 
195
 
 
 
196
  ## License and Responsible Use
197
 
 
198
  The source code is licensed under the [Apache License, Version 2.0](https://github.com/breezeblue-ai/breeze-tts/blob/main/LICENSE). The audio tokenizer is based on [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) by the Alibaba Qwen Team and is licensed under the Apache License, Version 2.0. Model weights, checkpoints, adapters, derivative models, and self-hosted outputs are governed separately by the [BreezeBlue Research and Non-Commercial License](https://huggingface.co/BreezeBlue/Breeze-TTS-2/blob/main/LICENSE). The Apache License does not grant rights to use the model commercially.
199
 
 
200
  If you have an active paid subscription, outputs you generate through BreezeBlue's hosted platform or API at [breezeblue.ai](https://breezeblue.ai/) can be used commercially, subject to our [Terms of Service](https://breezeblue.ai/legal/terms). A paid subscription does not grant commercial rights to the open-weight model or self-hosted outputs.
201
 
 
202
  You are responsible for complying with applicable laws and obtaining all necessary rights and consents for inputs, reference audio, voices, and outputs. Unauthorized voice cloning, impersonation, fraud, and other unlawful or harmful uses are prohibited.
203
 
 
204
  The code and Model Materials are provided "AS IS," without warranties or liability to the maximum extent permitted by law. Third-party components remain subject to their respective licenses.