ZichenAI commited on
Commit
a427608
Β·
verified Β·
1 Parent(s): 8449246

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +62 -4
README.md CHANGED
@@ -184,7 +184,61 @@ Caveats, stated plainly:
184
  - The Hub checkpoint stores float32 weights and the harness ran it at `bfloat16` (see Β§5.2).
185
  The two agree to within 0.002 on every task listed, so precision is not driving the table.
186
 
187
- ### 4.4 Effect of the P0 revision
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
188
 
189
  Measured with the same data, recipe and seed, P0 vs the old single-vector shared branch:
190
 
@@ -278,11 +332,15 @@ above, loads them at **1.20 GB** and leaves the (float32) rotary tables untouche
278
  (reuse the earlier rule/format, answer directly), not new reasoning ability.
279
  2. **The shared branch is still barely used** (gate β‰ˆ 0.0101 after SFT, same as the old
280
  revision). Whether P0 pays off can only be decided in a long-context / window-shrink regime.
281
- 3. **Trained on synthetic dialogues.** The 1,294 conversations are model-generated; style and
 
 
 
 
282
  coverage are limited to the six transfer kinds and the topics they cover. Evaluation above
283
  is on a held-out slice of the same distribution β€” not a general capability claim.
284
- 4. **Custom architecture, no llama.cpp/GGUF support** (SSA attention is not implemented there).
285
- 5. **Not an instruct model at large.** It is a 0.6B research model; expect terse answers and
286
  occasional arithmetic slips.
287
 
288
  ## 7. License
 
184
  - The Hub checkpoint stores float32 weights and the harness ran it at `bfloat16` (see Β§5.2).
185
  The two agree to within 0.002 on every task listed, so precision is not driving the table.
186
 
187
+ ### 4.4 Inference cost and speed (RTX 4090D, bfloat16)
188
+
189
+ Measured with the project's own `scripts/bench_resources.py` on the card that trained the model
190
+ (RTX 4090D 24 GB, no other process on it), bfloat16, `torch.no_grad()`, greedy decoding of 64
191
+ tokens after a prefill of 1,024 / 2,048 / 4,096 tokens, identical script for both models.
192
+
193
+ **The checkpoint exactly as shipped** (`ssa_force_full_window: true` β€” see below), 1,024-token
194
+ prefill:
195
+
196
+ | Metric | Qwen3-0.6B-Base | BaiHu-V1-Flash |
197
+ |---|---|---|
198
+ | Parameters (M) | 596.0 | 598.8 |
199
+ | Peak memory, prefill (GB) | 1.63 | **1.52** |
200
+ | Peak memory, generate (GB) | 1.78 | 2.01 |
201
+ | Prefill latency = TTFT (s) | 0.025 | 1.48 |
202
+ | TPOT β€” time per output token (ms) | 17.3 | 68.4 |
203
+ | Decode throughput (tok/s) | 57.8 | 14.6 |
204
+ | Attention FLOPs per token (GFLOPs) | 0.118 | 0.202 |
205
+ | Attention keys read per query (vs full attention) | 1.00Γ— | 1.72Γ— |
206
+ | GPU utilization mean (%) / power mean (W) | 15.5 / 69.8 | 20.0 / 69.0 |
207
+
208
+ How both quantities scale with context (same run, every cell in the order *base β†’ ours*):
209
+
210
+ | Prefill | TTFT (ms) | TPOT (ms) | Decode (tok/s) | Keys read per query | Generate peak (GB) |
211
+ |---|---|---|---|---|---|
212
+ | 1,024 | 25 β†’ 1,477 | 17.3 β†’ 68.4 | 57.8 β†’ 14.6 | 1.00Γ— β†’ 1.72Γ— | 1.78 β†’ 2.01 |
213
+ | 2,048 | 39 β†’ 2,113 | 19.8 β†’ 78.7 | 50.4 β†’ 12.7 | 1.00Γ— β†’ 1.45Γ— | 2.35 β†’ 2.82 |
214
+ | 4,096 | 72 β†’ 3,610 | 21.4 β†’ 108.8 | 46.7 β†’ 9.2 | 1.00Γ— β†’ 1.24Γ— | 3.53 β†’ 4.43 |
215
+
216
+ Stated plainly:
217
+
218
+ - **There is no speed advantage at any length tested.** Prefill (TTFT) is 50–59Γ— slower in
219
+ wall-clock β€” 1.48 s vs 25 ms at 1 K β€” decode is 3.9Γ— slower at 1 K and 5.1Γ— at 4 K, and peak
220
+ memory is comparable (slightly *lower* for this model during prefill, slightly higher during
221
+ generation).
222
+ - **The sparse path is not actually saving anything here.** This checkpoint is trained *and
223
+ released* with `ssa_force_full_window: true`: the dense local window always covers the entire
224
+ causal prefix, so the shared and sparse paths are **additive on top of full attention** instead
225
+ of replacing part of it. Hence keys read per query above 1.00Γ—.
226
+ - **Read the 1.72Γ— honestly.** It is the measured cost of the shipped configuration, and it is
227
+ also why the model spends more attention FLOPs per token than the dense base (0.202 vs 0.118
228
+ GFLOPs at 1 K) while still being far slower in wall-clock. Sparsity only starts to bite once
229
+ `top_k Γ— block_size` is small relative to the context (see the next point).
230
+ - **Turning the window shrinkage on is a flag, not a retrain** β€” but it was never validated in
231
+ that mode, because the weights were trained with the window forced open. For reference, with
232
+ `ssa_force_full_window=false` the same weights read **0.90Γ— / 0.54Γ— / 0.29Γ—** as many keys at
233
+ 1 K / 2 K / 4 K and decode at 20.7 / 18.5 / 14.8 tok/s. Even then it stays **2.5–3.2Γ— slower**
234
+ than the dense base, and quality in that mode is **unevaluated** β€” treat the numbers as the
235
+ architecture's ceiling on this implementation, not as a free win.
236
+ - What follows from this is an implementation item, not an architecture one: the per-block Python
237
+ loop in the SSA layer issues many small kernels, so launch overhead dominates the arithmetic
238
+ saved (GPU utilization never exceeds ~28%). The sparse path has to be fused before any of this
239
+ can pay off on real hardware.
240
+
241
+ ### 4.5 Effect of the P0 revision
242
 
243
  Measured with the same data, recipe and seed, P0 vs the old single-vector shared branch:
244
 
 
332
  (reuse the earlier rule/format, answer directly), not new reasoning ability.
333
  2. **The shared branch is still barely used** (gate β‰ˆ 0.0101 after SFT, same as the old
334
  revision). Whether P0 pays off can only be decided in a long-context / window-shrink regime.
335
+ 3. **No inference speedup β€” in fact a slowdown.** As shipped, the model is 3.9–5.1Γ— slower to
336
+ decode and ~50Γ— slower to prefill than the dense base, and reads *more* attention keys,
337
+ because the released configuration forces the dense window open (see Β§4.4). Use it to study
338
+ the architecture, not to serve traffic.
339
+ 4. **Trained on synthetic dialogues.** The 1,294 conversations are model-generated; style and
340
  coverage are limited to the six transfer kinds and the topics they cover. Evaluation above
341
  is on a held-out slice of the same distribution β€” not a general capability claim.
342
+ 5. **Custom architecture, no llama.cpp/GGUF support** (SSA attention is not implemented there).
343
+ 6. **Not an instruct model at large.** It is a 0.6B research model; expect terse answers and
344
  occasional arithmetic slips.
345
 
346
  ## 7. License