Gemma 4 26B-A4B IT FP8 Dynamic Norouter

Production-ready offline FP8 checkpoint for vLLM โ€” 47% less VRAM, 80% more concurrency vs BF16.

We searched for a usable offline FP8 checkpoint of Gemma 4 26B-A4B-it but couldn't find one that worked cleanly with vLLM. So we vibe-coded our own and are sharing it with the community.

This repository hosts an offline FP8 checkpoint derived from google/gemma-4-26B-A4B-it for vLLM serving. No on-the-fly quantization needed at startup.

Published by Largitdata Inc.

Note: This is a derived operational checkpoint, not an official Google release. The original model's license terms, safety guidance, and documentation remain authoritative.

๐Ÿ“– ไธญๆ–‡่ชชๆ˜Ž / Chinese Version


Model Details

  • Base model: google/gemma-4-26B-A4B-it
  • Derived format: offline FP8 checkpoint for vLLM
  • Quantization tool: llmcompressor
  • Quantization method: FP8_DYNAMIC
  • Calibration data: None required (dynamic quantization)
  • Excluded weights:
    • norm-class 1D tensors โ€” excluded to avoid expected 2D linear weight validation errors during quantization
    • re:.*router\.proj$ โ€” MoE router weights excluded to maintain compatibility with the Gemma4 vLLM loading path
  • Output directory name: gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER
  • Primary serving target: vllm/vllm-openai:gemma4
  • Organization: Largitdata Inc.

Test Environment

  • GPU: NVIDIA H200 NVL (143 GB VRAM)
  • Runtime: vllm/vllm-openai:gemma4
  • KV cache dtype: fp8
  • max_model_len: 32768
  • gpu_memory_utilization: 0.55

Observed vLLM startup characteristics:

  • model weight loading: 15.76 s
  • model loading total: 16.88 s
  • torch.compile: 56.98 s
  • engine init: 102.17 s
  • total time to /v1/models ready: about 153 s

Observed runtime capacity:

  • max_num_batched_tokens = 8192
  • available KV cache memory: 46.37 GiB
  • GPU KV cache size: 405,184 tokens
  • maximum concurrency at 32,768 tokens/request: 38.87x

Serving Capacity Comparison

Metric FP8 Dynamic Norouter BF16 Baseline
Model loading memory 25.75 GiB 48.5 GiB
GPU KV cache size 405,184 tokens 225,376 tokens
Max concurrency @ 32K tokens/req 38.87x 21.62x
VRAM savings 47% less โ€”
KV cache gain 80% more โ€”

Basic Benchmark

Single-request warm benchmark against the OpenAI-compatible vLLM endpoint:

  • prompt tokens: 38
  • completion tokens: 256
  • temperature: 0
Metric FP8 Dynamic Norouter BF16 Baseline
Avg end-to-end latency 1.629 s 1.536 s
Avg completion throughput 157.19 tok/s 166.62 tok/s
Avg total throughput 180.53 tok/s 191.36 tok/s

These numbers are single-request warm-path measurements, not multi-client throughput tests. In production multi-client scenarios, the FP8 variant's larger KV cache is expected to provide superior aggregate throughput.

BF16 is ~6% faster on single-request latency, but the FP8 variant uses 47% less VRAM and provides 80% more KV cache capacity. For production environments serving multiple concurrent users, the FP8 variant offers a better trade-off.

Accuracy Evaluation โ€” MMLU-Pro

We ran MMLU-Pro (full set, 14 categories) comparing this FP8 checkpoint against its BF16 baseline on the same harness and prompt template.

Model MMLU-Pro Accuracy Wall clock
26B BF16 (baseline) 81.59% 65 min
26B FP8 Dynamic Norouter (this) 81.33% 63 min
ฮ” (FP8 โˆ’ BF16) โˆ’0.26 pp โ€”

FP8 quantization cost is essentially zero โ€” the 0.26 pp gap sits inside benchmark noise. Our 26B BF16 number (81.59%) is ~1.0 pp below Google's reported 82.6%, consistent with harness/prompt differences rather than a quantization defect; both BF16 and FP8 shift together.

For context, we also evaluated the sibling 31B variant under identical conditions:

Model MMLU-Pro
26B BF16 81.59%
26B FP8 81.33%
31B BF16 84.60%
31B FP8 Dynamic 84.55%

See also: LargitData/gemma-4-31b-it-fp8.

Usage

Example vLLM launch:

docker run -d \
  --name vllm-gemma4-26b-fp8-norouter \
  --restart unless-stopped \
  --ipc=host \
  --shm-size 16G \
  --gpus all \
  -v /models \
  -p 8001:8000 \
  -e NVIDIA_VISIBLE_DEVICES=0 \
  vllm/vllm-openai:gemma4 \
  --model /models/gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER \
  --trust-remote-code \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.55 \
  --max-model-len 32768 \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --host 0.0.0.0 \
  --port 8000

Known Limitations

  • Single-request latency is ~6% higher than BF16 due to FP8 dequantization overhead.
  • MMLU-Pro measured (see above); MT-Bench and other benchmarks not yet run. Community contributions are welcome.
  • Only tested on NVIDIA H200 NVL. Other GPUs (A100, H100) may require adjusting gpu-memory-utilization.
  • MoE router weights (router.proj) and norm-class 1D tensors are excluded from quantization for vLLM compatibility. No routing degradation has been observed, but systematic evaluation has not been performed.

Intended Use

This artifact is intended for:

  • Operational vLLM deployment on H200-class hardware
  • Reproducible offline FP8 serving experiments
  • Environments where startup-time on-the-fly quantization is undesirable
  • Production inference with higher concurrency requirements

This artifact is not intended to replace the original base model documentation, safety guidance, or license terms.

License

This repository contains a derived checkpoint based on google/gemma-4-26B-A4B-it. Usage is subject to the Gemma Terms of Use.

Citation

If you use this artifact, please cite both the derived checkpoint and the upstream base model.

@misc{largitdata_gemma4_26b_a4b_it_fp8_dynamic_norouter_2026,
  title        = {Gemma 4 26B-A4B IT FP8 Dynamic Norouter},
  author       = {David Chiu},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/largitdata-inc/gemma-4-26b-a4b-it-fp8-dynamic-norouter}},
  note         = {Derived offline FP8 checkpoint from google/gemma-4-26B-A4B-it for vLLM serving, published by Largitdata Inc. \url{https://www.largitdata.com/}}
}

@misc{google_gemma4_26b_a4b_it,
  title        = {Gemma 4 26B-A4B IT},
  author       = {Google},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/google/gemma-4-26B-A4B-it}}
}

Disclaimer

Users are responsible for verifying license compatibility, downstream serving behavior, numerical quality, and safety characteristics for their own environment.


ไธญๆ–‡่ชชๆ˜Ž

็ตฆ vLLM ็”จ็š„้›ข็ทš FP8 checkpoint โ€” ๆฏ” BF16 ็œ 47% VRAM๏ผŒๅนณ่กŒ่™•็†่ƒฝๅŠ›ๅคš 80%ใ€‚

ๆˆ‘ๅ€‘ๅœจ็ถฒ่ทฏไธŠๆ‰พไบ†ไธ€่ผช๏ผŒๆฒ’ๆœ‰ๆ‰พๅˆฐๅ ช็”จ็š„ Gemma 4 26B ้›ข็ทš FP8 ็‰ˆๆœฌ๏ผŒ็ดขๆ€ง่‡ชๅทฑ vibe coding ๅšไบ†ไธ€็‰ˆ๏ผŒ่ฒข็ป็ตฆ็คพ็พคใ€‚

้€™ๅ€‹ Repo ๆไพ›ๅพž google/gemma-4-26B-A4B-it ่ก็”Ÿๅ‡บ็š„้›ข็ทš FP8 checkpoint๏ผŒ่ฎ“ vLLM ๅฏไปฅ็›ดๆŽฅ่ผ‰ๅ…ฅๆœๅ‹™๏ผŒไธ้œ€่ฆๅœจๅ•Ÿๅ‹•ๆ™‚ๅŸท่กŒ on-the-fly ้‡ๅŒ–ใ€‚

็”ฑ Largitdata Inc. ็™ผไฝˆใ€‚

ๆณจๆ„๏ผš ้€™ๆ˜ฏ่ก็”Ÿ็š„ๆ“ไฝœ็”จ checkpoint๏ผŒไธฆ้ž Google ๅฎ˜ๆ–น็™ผไฝˆใ€‚ๅŽŸๅง‹ๆจกๅž‹็š„ๆŽˆๆฌŠๆขๆฌพใ€ๅฎ‰ๅ…จๆŒ‡ๅผ•่ˆ‡ๆ–‡ไปถไปไปฅๅฎ˜ๆ–น็‚บๆบ–ใ€‚

ๆจกๅž‹็ดฐ็ฏ€

  • ๅŸบๅบ•ๆจกๅž‹๏ผš google/gemma-4-26B-A4B-it
  • ๆ ผๅผ๏ผš ้›ข็ทš FP8 checkpoint๏ผŒไพ› vLLM ไฝฟ็”จ
  • ้‡ๅŒ–ๅทฅๅ…ท๏ผš llmcompressor
  • ้‡ๅŒ–ๆ–นๅผ๏ผš FP8_DYNAMIC
  • ๆ กๆบ–่ณ‡ๆ–™๏ผš ไธ้œ€่ฆ๏ผˆๅ‹•ๆ…‹้‡ๅŒ–๏ผ‰
  • ๆŽ’้™ค็š„ๆฌŠ้‡๏ผš
    • norm ้กžไธ€็ถญ tensor โ€” ้ฟๅ…้‡ๅŒ–้ฉ—่ญ‰ๆ™‚็”ข็”Ÿ expected 2D linear weight ้กž้Œฏ่ชค
    • re:.*router\.proj$ โ€” MoE router ๆฌŠ้‡๏ผŒ็ถญๆŒ่ˆ‡ Gemma4 vLLM ่ผ‰ๅ…ฅ่ทฏๅพ‘็š„็›ธๅฎนๆ€ง
  • ไธป่ฆ้ƒจ็ฝฒ็›ฎๆจ™๏ผš vllm/vllm-openai:gemma4

ๆธฌ่ฉฆ็’ฐๅขƒ

  • GPU๏ผš NVIDIA H200 NVL๏ผˆ143 GB VRAM๏ผ‰
  • Runtime๏ผš vllm/vllm-openai:gemma4
  • KV cache dtype๏ผš fp8
  • max_model_len๏ผš 32768
  • gpu_memory_utilization๏ผš 0.55

ๅ•Ÿๅ‹•ๅฏฆๆธฌๆ•ธๆ“š๏ผš

  • ๆจกๅž‹ๆฌŠ้‡่ผ‰ๅ…ฅ๏ผš15.76 s
  • ๆจกๅž‹่ผ‰ๅ…ฅ็ธฝ่จˆ๏ผš16.88 s
  • torch.compile๏ผš56.98 s
  • ๅผ•ๆ“Žๅˆๅง‹ๅŒ–๏ผš102.17 s
  • /v1/models ๅฐฑ็ท’็ธฝๆ™‚้–“๏ผš็ด„ 153 s

ๅŸท่กŒๆœŸๅฎน้‡๏ผš

  • max_num_batched_tokens = 8192
  • ๅฏ็”จ KV cache ่จ˜ๆ†ถ้ซ”๏ผš46.37 GiB
  • GPU KV cache ๅคงๅฐ๏ผš405,184 tokens
  • ๆœ€ๅคงๅนณ่กŒ่™•็†้‡๏ผˆ32,768 tokens/request๏ผ‰๏ผš38.87x

ๆœๅ‹™ๅฎน้‡ๆฏ”่ผƒ

ๆŒ‡ๆจ™ FP8 Dynamic Norouter BF16 ๅŽŸ็‰ˆ
ๆจกๅž‹่ผ‰ๅ…ฅ่จ˜ๆ†ถ้ซ” 25.75 GiB 48.5 GiB
GPU KV cache ๅคงๅฐ 405,184 tokens 225,376 tokens
ๆœ€ๅคงๅนณ่กŒ่™•็†้‡ @ 32K tokens/req 38.87x 21.62x
VRAM ็ฏ€็œ 47% โ€”
KV cache ๅขžๅŠ  80% โ€”

ๅŸบ็คŽๆ•ˆ่ƒฝๆธฌ่ฉฆ

ๅ–ฎ่ซ‹ๆฑ‚ๆš–ๆฉŸๆธฌ่ฉฆ๏ผˆOpenAI ็›ธๅฎน vLLM endpoint๏ผ‰๏ผš

  • prompt tokens๏ผš38
  • completion tokens๏ผš256
  • temperature๏ผš0
ๆŒ‡ๆจ™ FP8 Dynamic Norouter BF16 ๅŽŸ็‰ˆ
ๅนณๅ‡็ซฏๅˆฐ็ซฏๅปถ้ฒ 1.629 s 1.536 s
ๅนณๅ‡ completion ๅžๅ้‡ 157.19 tok/s 166.62 tok/s
ๅนณๅ‡็ธฝๅžๅ้‡ 180.53 tok/s 191.36 tok/s

ไปฅไธŠ็‚บๅ–ฎ่ซ‹ๆฑ‚ๆš–ๆฉŸ่ทฏๅพ‘ๆธฌ้‡ๅ€ผ๏ผŒ้žๅคš็”จๆˆถๅžๅ้‡ๆธฌ่ฉฆใ€‚ๅœจ็”Ÿ็”ข็’ฐๅขƒๅคš็”จๆˆถๅ ดๆ™ฏไธ‹๏ผŒFP8 ็‰ˆๆœฌๆ›ดๅคง็š„ KV cache ้ ๆœŸ่ƒฝๆไพ›ๆ›ดๅฅฝ็š„ๆ•ด้ซ”ๅžๅ้‡ใ€‚

็ต่ซ–๏ผšBF16 ๅ–ฎ่ซ‹ๆฑ‚็•ฅๅฟซ๏ผˆ็ด„ 6%๏ผ‰๏ผŒไฝ† FP8 ็‰ˆๆœฌ VRAM ็”จ้‡ๆธ›ๅฐ‘ 47%๏ผŒๅฏ็”จ KV cache ๅขžๅŠ  80%ใ€‚้œ€่ฆๅŒๆ™‚ๆœๅ‹™ๅคš็”จๆˆถ็š„็”Ÿ็”ข็’ฐๅขƒ๏ผŒFP8 ็‰ˆๆœฌๆ›ดๅ…ทๅ„ชๅ‹ขใ€‚

็ฒพๅบฆ่ฉ•ไผฐ โ€” MMLU-Pro

ไฝฟ็”จ็›ธๅŒ harness ่ˆ‡ prompt template๏ผŒๅฐ FP8 checkpoint ่ˆ‡ BF16 ๅŽŸ็‰ˆๅŒๆ™‚ๅŸท่กŒ MMLU-Pro๏ผˆๅ…จ้›†๏ผŒ14 ้กžๅˆฅ๏ผ‰๏ผš

ๆจกๅž‹ MMLU-Pro ๆบ–็ขบ็އ ็ธฝๅŸท่กŒๆ™‚้–“
26B BF16๏ผˆๅŸบๆบ–๏ผ‰ 81.59% 65 min
26B FP8 Dynamic Norouter๏ผˆๆœฌ็‰ˆๆœฌ๏ผ‰ 81.33% 63 min
ฮ” (FP8 โˆ’ BF16) โˆ’0.26 pp โ€”

FP8 ้‡ๅŒ–็š„ไปฃๅƒนๅŸบๆœฌ็‚บ้›ถ โ€” 0.26 pp ็š„ๅทฎ่ท่ฝๅœจ benchmark ๅ™ช้Ÿณ็ฏ„ๅœๅ…งใ€‚ๆˆ‘ๅ€‘็š„ 26B BF16 ๆˆ็ธพ๏ผˆ81.59%๏ผ‰ๆฏ” Google ๅฎ˜ๆ–น็š„ 82.6% ไฝŽ็ด„ 1.0 pp๏ผŒ้€™ๆ˜ฏ harness / prompt ๅทฎ็•ฐ้€ ๆˆ๏ผŒ้ž้‡ๅŒ–ๅ•้กŒ๏ผ›BF16 ่ˆ‡ FP8 ๅ…ฉ่€…ไธ€่ตทๅ็งปใ€‚

ไฝœ็‚บๅฐ็…ง๏ผŒๅŒๆขไปถไธ‹ 31B ็‰ˆๆœฌ็ตๆžœ๏ผš

ๆจกๅž‹ MMLU-Pro
26B BF16 81.59%
26B FP8 81.33%
31B BF16 84.60%
31B FP8 Dynamic 84.55%

ๅฆ่ฆ‹๏ผšLargitData/gemma-4-31b-it-fp8ใ€‚

ไฝฟ็”จๆ–นๅผ

vLLM ๅ•Ÿๅ‹•็ฏ„ไพ‹๏ผš

docker run -d \
  --name vllm-gemma4-26b-fp8-norouter \
  --restart unless-stopped \
  --ipc=host \
  --shm-size 16G \
  --gpus all \
  -v /models \
  -p 8001:8000 \
  -e NVIDIA_VISIBLE_DEVICES=0 \
  vllm/vllm-openai:gemma4 \
  --model /models/gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER \
  --trust-remote-code \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.55 \
  --max-model-len 32768 \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --host 0.0.0.0 \
  --port 8000

ๅทฒ็Ÿฅ้™ๅˆถ

  • ๅ–ฎ่ซ‹ๆฑ‚ๅปถ้ฒๆฏ” BF16 ้ซ˜็ด„ 6%๏ผŒไธปๅ› ็‚บ FP8 dequantization ็š„้กๅค–้–‹้Šท
  • MMLU-Pro ๅทฒๅฎŒๆˆ๏ผˆ่ฆ‹ไธŠ๏ผ‰๏ผŒMT-Bench ็ญ‰ๅ…ถไป– benchmark ๅฐšๆœชๅŸท่กŒ๏ผŒๆญก่ฟŽ็คพ็พค่ฃœๅ……
  • ๅƒ…ๅœจ H200 NVL ไธŠๅฏฆๆธฌ๏ผŒๅ…ถไป– GPU๏ผˆๅฆ‚ A100ใ€H100๏ผ‰ๅฏ่ƒฝ้œ€่ฆ่ชฟๆ•ด gpu-memory-utilization
  • MoE router ๆฌŠ้‡๏ผˆrouter.proj๏ผ‰่ˆ‡ norm ้กžไธ€็ถญ tensor ่ขซๆŽ’้™คๅœจ้‡ๅŒ–็ฏ„ๅœๅค–ไปฅ็ถญๆŒ vLLM ็›ธๅฎนๆ€ง๏ผŒ็›ฎๅ‰ๆœช่ง€ๅฏŸๅˆฐๅˆ†ๆตๅ“่ณชไธ‹้™๏ผŒไฝ†ๅฐš็„ก็ณป็ตฑๆ€ง่ฉ•ไผฐ

ไฝฟ็”จๅ ดๆ™ฏ

ๆญค checkpoint ้ฉ็”จๆ–ผ๏ผš

  • ๅœจ H200 ็ญ‰็ดš็กฌ้ซ”ไธŠไปฅ vLLM ้€ฒ่กŒ็”Ÿ็”ข้ƒจ็ฝฒ
  • ๅฏ้‡็พ็š„้›ข็ทš FP8 ๆœๅ‹™ๅฏฆ้ฉ—
  • ไธๅธŒๆœ›ๅœจๅ•Ÿๅ‹•ๆ™‚ๅŸท่กŒ on-the-fly ้‡ๅŒ–็š„็’ฐๅขƒ
  • ้œ€่ฆๆ›ด้ซ˜ๅนณ่กŒ่™•็†่ƒฝๅŠ›็š„็”Ÿ็”ขๆŽจ่ซ–ๅ ดๆ™ฏ

ๆญค checkpoint ไธๅ–ไปฃๅŽŸๅง‹ๅŸบๅบ•ๆจกๅž‹็š„ๆ–‡ไปถใ€ๅฎ‰ๅ…จๆŒ‡ๅผ•ๆˆ–ๆŽˆๆฌŠๆขๆฌพใ€‚

ๆŽˆๆฌŠ

ๆญค Repo ๅŒ…ๅซๅŸบๆ–ผ google/gemma-4-26B-A4B-it ็š„่ก็”Ÿ checkpoint๏ผŒไฝฟ็”จ้ ˆ้ตๅฎˆ Gemma ไฝฟ็”จๆขๆฌพใ€‚

ๅ…่ฒฌ่ฒๆ˜Ž

ไฝฟ็”จ่€…้œ€่‡ช่กŒ้ฉ—่ญ‰ๆŽˆๆฌŠ็›ธๅฎนๆ€งใ€ไธ‹ๆธธๆœๅ‹™่กŒ็‚บใ€ๆ•ธๅ€ผๅ“่ณช่ˆ‡ๅฎ‰ๅ…จ็‰นๆ€งใ€‚

Downloads last month
34,254
Safetensors
Model size
26B params
Tensor type
BF16
ยท
F8_E4M3
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LargitData/gemma-4-26b-a4b-it-fp8

Quantized
(379)
this model

Space using LargitData/gemma-4-26b-a4b-it-fp8 1