Text Generation
Transformers
Safetensors
Chinese
English
baihu_ssa
sparse-attention
subq
ssa
long-context
supervised-fine-tuning
transfer-learning
commercial-license-required
conversational
Instructions to use ZichenAI/BaiHu-V1-Flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ZichenAI/BaiHu-V1-Flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ZichenAI/BaiHu-V1-Flash") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("ZichenAI/BaiHu-V1-Flash", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ZichenAI/BaiHu-V1-Flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZichenAI/BaiHu-V1-Flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZichenAI/BaiHu-V1-Flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZichenAI/BaiHu-V1-Flash
- SGLang
How to use ZichenAI/BaiHu-V1-Flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ZichenAI/BaiHu-V1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZichenAI/BaiHu-V1-Flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ZichenAI/BaiHu-V1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZichenAI/BaiHu-V1-Flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ZichenAI/BaiHu-V1-Flash with Docker Model Runner:
docker model run hf.co/ZichenAI/BaiHu-V1-Flash
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -184,7 +184,61 @@ Caveats, stated plainly:
|
|
| 184 |
- The Hub checkpoint stores float32 weights and the harness ran it at `bfloat16` (see Β§5.2).
|
| 185 |
The two agree to within 0.002 on every task listed, so precision is not driving the table.
|
| 186 |
|
| 187 |
-
### 4.4
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 188 |
|
| 189 |
Measured with the same data, recipe and seed, P0 vs the old single-vector shared branch:
|
| 190 |
|
|
@@ -278,11 +332,15 @@ above, loads them at **1.20 GB** and leaves the (float32) rotary tables untouche
|
|
| 278 |
(reuse the earlier rule/format, answer directly), not new reasoning ability.
|
| 279 |
2. **The shared branch is still barely used** (gate β 0.0101 after SFT, same as the old
|
| 280 |
revision). Whether P0 pays off can only be decided in a long-context / window-shrink regime.
|
| 281 |
-
3. **
|
|
|
|
|
|
|
|
|
|
|
|
|
| 282 |
coverage are limited to the six transfer kinds and the topics they cover. Evaluation above
|
| 283 |
is on a held-out slice of the same distribution β not a general capability claim.
|
| 284 |
-
|
| 285 |
-
|
| 286 |
occasional arithmetic slips.
|
| 287 |
|
| 288 |
## 7. License
|
|
|
|
| 184 |
- The Hub checkpoint stores float32 weights and the harness ran it at `bfloat16` (see Β§5.2).
|
| 185 |
The two agree to within 0.002 on every task listed, so precision is not driving the table.
|
| 186 |
|
| 187 |
+
### 4.4 Inference cost and speed (RTX 4090D, bfloat16)
|
| 188 |
+
|
| 189 |
+
Measured with the project's own `scripts/bench_resources.py` on the card that trained the model
|
| 190 |
+
(RTX 4090D 24 GB, no other process on it), bfloat16, `torch.no_grad()`, greedy decoding of 64
|
| 191 |
+
tokens after a prefill of 1,024 / 2,048 / 4,096 tokens, identical script for both models.
|
| 192 |
+
|
| 193 |
+
**The checkpoint exactly as shipped** (`ssa_force_full_window: true` β see below), 1,024-token
|
| 194 |
+
prefill:
|
| 195 |
+
|
| 196 |
+
| Metric | Qwen3-0.6B-Base | BaiHu-V1-Flash |
|
| 197 |
+
|---|---|---|
|
| 198 |
+
| Parameters (M) | 596.0 | 598.8 |
|
| 199 |
+
| Peak memory, prefill (GB) | 1.63 | **1.52** |
|
| 200 |
+
| Peak memory, generate (GB) | 1.78 | 2.01 |
|
| 201 |
+
| Prefill latency = TTFT (s) | 0.025 | 1.48 |
|
| 202 |
+
| TPOT β time per output token (ms) | 17.3 | 68.4 |
|
| 203 |
+
| Decode throughput (tok/s) | 57.8 | 14.6 |
|
| 204 |
+
| Attention FLOPs per token (GFLOPs) | 0.118 | 0.202 |
|
| 205 |
+
| Attention keys read per query (vs full attention) | 1.00Γ | 1.72Γ |
|
| 206 |
+
| GPU utilization mean (%) / power mean (W) | 15.5 / 69.8 | 20.0 / 69.0 |
|
| 207 |
+
|
| 208 |
+
How both quantities scale with context (same run, every cell in the order *base β ours*):
|
| 209 |
+
|
| 210 |
+
| Prefill | TTFT (ms) | TPOT (ms) | Decode (tok/s) | Keys read per query | Generate peak (GB) |
|
| 211 |
+
|---|---|---|---|---|---|
|
| 212 |
+
| 1,024 | 25 β 1,477 | 17.3 β 68.4 | 57.8 β 14.6 | 1.00Γ β 1.72Γ | 1.78 β 2.01 |
|
| 213 |
+
| 2,048 | 39 β 2,113 | 19.8 β 78.7 | 50.4 β 12.7 | 1.00Γ β 1.45Γ | 2.35 β 2.82 |
|
| 214 |
+
| 4,096 | 72 β 3,610 | 21.4 β 108.8 | 46.7 β 9.2 | 1.00Γ β 1.24Γ | 3.53 β 4.43 |
|
| 215 |
+
|
| 216 |
+
Stated plainly:
|
| 217 |
+
|
| 218 |
+
- **There is no speed advantage at any length tested.** Prefill (TTFT) is 50β59Γ slower in
|
| 219 |
+
wall-clock β 1.48 s vs 25 ms at 1 K β decode is 3.9Γ slower at 1 K and 5.1Γ at 4 K, and peak
|
| 220 |
+
memory is comparable (slightly *lower* for this model during prefill, slightly higher during
|
| 221 |
+
generation).
|
| 222 |
+
- **The sparse path is not actually saving anything here.** This checkpoint is trained *and
|
| 223 |
+
released* with `ssa_force_full_window: true`: the dense local window always covers the entire
|
| 224 |
+
causal prefix, so the shared and sparse paths are **additive on top of full attention** instead
|
| 225 |
+
of replacing part of it. Hence keys read per query above 1.00Γ.
|
| 226 |
+
- **Read the 1.72Γ honestly.** It is the measured cost of the shipped configuration, and it is
|
| 227 |
+
also why the model spends more attention FLOPs per token than the dense base (0.202 vs 0.118
|
| 228 |
+
GFLOPs at 1 K) while still being far slower in wall-clock. Sparsity only starts to bite once
|
| 229 |
+
`top_k Γ block_size` is small relative to the context (see the next point).
|
| 230 |
+
- **Turning the window shrinkage on is a flag, not a retrain** β but it was never validated in
|
| 231 |
+
that mode, because the weights were trained with the window forced open. For reference, with
|
| 232 |
+
`ssa_force_full_window=false` the same weights read **0.90Γ / 0.54Γ / 0.29Γ** as many keys at
|
| 233 |
+
1 K / 2 K / 4 K and decode at 20.7 / 18.5 / 14.8 tok/s. Even then it stays **2.5β3.2Γ slower**
|
| 234 |
+
than the dense base, and quality in that mode is **unevaluated** β treat the numbers as the
|
| 235 |
+
architecture's ceiling on this implementation, not as a free win.
|
| 236 |
+
- What follows from this is an implementation item, not an architecture one: the per-block Python
|
| 237 |
+
loop in the SSA layer issues many small kernels, so launch overhead dominates the arithmetic
|
| 238 |
+
saved (GPU utilization never exceeds ~28%). The sparse path has to be fused before any of this
|
| 239 |
+
can pay off on real hardware.
|
| 240 |
+
|
| 241 |
+
### 4.5 Effect of the P0 revision
|
| 242 |
|
| 243 |
Measured with the same data, recipe and seed, P0 vs the old single-vector shared branch:
|
| 244 |
|
|
|
|
| 332 |
(reuse the earlier rule/format, answer directly), not new reasoning ability.
|
| 333 |
2. **The shared branch is still barely used** (gate β 0.0101 after SFT, same as the old
|
| 334 |
revision). Whether P0 pays off can only be decided in a long-context / window-shrink regime.
|
| 335 |
+
3. **No inference speedup β in fact a slowdown.** As shipped, the model is 3.9β5.1Γ slower to
|
| 336 |
+
decode and ~50Γ slower to prefill than the dense base, and reads *more* attention keys,
|
| 337 |
+
because the released configuration forces the dense window open (see Β§4.4). Use it to study
|
| 338 |
+
the architecture, not to serve traffic.
|
| 339 |
+
4. **Trained on synthetic dialogues.** The 1,294 conversations are model-generated; style and
|
| 340 |
coverage are limited to the six transfer kinds and the topics they cover. Evaluation above
|
| 341 |
is on a held-out slice of the same distribution β not a general capability claim.
|
| 342 |
+
5. **Custom architecture, no llama.cpp/GGUF support** (SSA attention is not implemented there).
|
| 343 |
+
6. **Not an instruct model at large.** It is a 0.6B research model; expect terse answers and
|
| 344 |
occasional arithmetic slips.
|
| 345 |
|
| 346 |
## 7. License
|