Instructions to use LargitData/gemma-4-26b-a4b-it-fp8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LargitData/gemma-4-26b-a4b-it-fp8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="LargitData/gemma-4-26b-a4b-it-fp8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("LargitData/gemma-4-26b-a4b-it-fp8") model = AutoModelForMultimodalLM.from_pretrained("LargitData/gemma-4-26b-a4b-it-fp8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LargitData/gemma-4-26b-a4b-it-fp8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LargitData/gemma-4-26b-a4b-it-fp8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LargitData/gemma-4-26b-a4b-it-fp8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/LargitData/gemma-4-26b-a4b-it-fp8
- SGLang
How to use LargitData/gemma-4-26b-a4b-it-fp8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LargitData/gemma-4-26b-a4b-it-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LargitData/gemma-4-26b-a4b-it-fp8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LargitData/gemma-4-26b-a4b-it-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LargitData/gemma-4-26b-a4b-it-fp8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use LargitData/gemma-4-26b-a4b-it-fp8 with Docker Model Runner:
docker model run hf.co/LargitData/gemma-4-26b-a4b-it-fp8
Gemma 4 26B-A4B IT FP8 Dynamic Norouter
Production-ready offline FP8 checkpoint for vLLM โ 47% less VRAM, 80% more concurrency vs BF16.
We searched for a usable offline FP8 checkpoint of Gemma 4 26B-A4B-it but couldn't find one that worked cleanly with vLLM. So we vibe-coded our own and are sharing it with the community.
This repository hosts an offline FP8 checkpoint derived from google/gemma-4-26B-A4B-it for vLLM serving. No on-the-fly quantization needed at startup.
Published by Largitdata Inc.
Note: This is a derived operational checkpoint, not an official Google release. The original model's license terms, safety guidance, and documentation remain authoritative.
๐ ไธญๆ่ชชๆ / Chinese Version
Model Details
- Base model:
google/gemma-4-26B-A4B-it - Derived format: offline FP8 checkpoint for vLLM
- Quantization tool:
llmcompressor - Quantization method:
FP8_DYNAMIC - Calibration data: None required (dynamic quantization)
- Excluded weights:
norm-class 1D tensors โ excluded to avoidexpected 2D linear weightvalidation errors during quantizationre:.*router\.proj$โ MoE router weights excluded to maintain compatibility with the Gemma4 vLLM loading path
- Output directory name:
gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER - Primary serving target:
vllm/vllm-openai:gemma4 - Organization: Largitdata Inc.
Test Environment
- GPU:
NVIDIA H200 NVL(143 GB VRAM) - Runtime:
vllm/vllm-openai:gemma4 - KV cache dtype:
fp8 max_model_len:32768gpu_memory_utilization:0.55
Observed vLLM startup characteristics:
- model weight loading:
15.76 s - model loading total:
16.88 s torch.compile:56.98 s- engine init:
102.17 s - total time to
/v1/modelsready: about153 s
Observed runtime capacity:
max_num_batched_tokens = 8192- available KV cache memory:
46.37 GiB - GPU KV cache size:
405,184 tokens - maximum concurrency at
32,768tokens/request:38.87x
Serving Capacity Comparison
| Metric | FP8 Dynamic Norouter | BF16 Baseline |
|---|---|---|
| Model loading memory | 25.75 GiB | 48.5 GiB |
| GPU KV cache size | 405,184 tokens | 225,376 tokens |
| Max concurrency @ 32K tokens/req | 38.87x | 21.62x |
| VRAM savings | 47% less | โ |
| KV cache gain | 80% more | โ |
Basic Benchmark
Single-request warm benchmark against the OpenAI-compatible vLLM endpoint:
- prompt tokens:
38 - completion tokens:
256 - temperature:
0
| Metric | FP8 Dynamic Norouter | BF16 Baseline |
|---|---|---|
| Avg end-to-end latency | 1.629 s | 1.536 s |
| Avg completion throughput | 157.19 tok/s | 166.62 tok/s |
| Avg total throughput | 180.53 tok/s | 191.36 tok/s |
These numbers are single-request warm-path measurements, not multi-client throughput tests. In production multi-client scenarios, the FP8 variant's larger KV cache is expected to provide superior aggregate throughput.
BF16 is ~6% faster on single-request latency, but the FP8 variant uses 47% less VRAM and provides 80% more KV cache capacity. For production environments serving multiple concurrent users, the FP8 variant offers a better trade-off.
Accuracy Evaluation โ MMLU-Pro
We ran MMLU-Pro (full set, 14 categories) comparing this FP8 checkpoint against its BF16 baseline on the same harness and prompt template.
| Model | MMLU-Pro Accuracy | Wall clock |
|---|---|---|
| 26B BF16 (baseline) | 81.59% | 65 min |
| 26B FP8 Dynamic Norouter (this) | 81.33% | 63 min |
| ฮ (FP8 โ BF16) | โ0.26 pp | โ |
FP8 quantization cost is essentially zero โ the 0.26 pp gap sits inside benchmark noise. Our 26B BF16 number (81.59%) is ~1.0 pp below Google's reported 82.6%, consistent with harness/prompt differences rather than a quantization defect; both BF16 and FP8 shift together.
For context, we also evaluated the sibling 31B variant under identical conditions:
| Model | MMLU-Pro |
|---|---|
| 26B BF16 | 81.59% |
| 26B FP8 | 81.33% |
| 31B BF16 | 84.60% |
| 31B FP8 Dynamic | 84.55% |
See also: LargitData/gemma-4-31b-it-fp8.
Usage
Example vLLM launch:
docker run -d \
--name vllm-gemma4-26b-fp8-norouter \
--restart unless-stopped \
--ipc=host \
--shm-size 16G \
--gpus all \
-v /models \
-p 8001:8000 \
-e NVIDIA_VISIBLE_DEVICES=0 \
vllm/vllm-openai:gemma4 \
--model /models/gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.55 \
--max-model-len 32768 \
--enable-auto-tool-choice \
--tool-call-parser gemma4 \
--host 0.0.0.0 \
--port 8000
Known Limitations
- Single-request latency is ~6% higher than BF16 due to FP8 dequantization overhead.
- MMLU-Pro measured (see above); MT-Bench and other benchmarks not yet run. Community contributions are welcome.
- Only tested on NVIDIA H200 NVL. Other GPUs (A100, H100) may require adjusting
gpu-memory-utilization. - MoE router weights (
router.proj) andnorm-class 1D tensors are excluded from quantization for vLLM compatibility. No routing degradation has been observed, but systematic evaluation has not been performed.
Intended Use
This artifact is intended for:
- Operational vLLM deployment on H200-class hardware
- Reproducible offline FP8 serving experiments
- Environments where startup-time on-the-fly quantization is undesirable
- Production inference with higher concurrency requirements
This artifact is not intended to replace the original base model documentation, safety guidance, or license terms.
License
This repository contains a derived checkpoint based on google/gemma-4-26B-A4B-it. Usage is subject to the Gemma Terms of Use.
Citation
If you use this artifact, please cite both the derived checkpoint and the upstream base model.
@misc{largitdata_gemma4_26b_a4b_it_fp8_dynamic_norouter_2026,
title = {Gemma 4 26B-A4B IT FP8 Dynamic Norouter},
author = {David Chiu},
year = {2026},
howpublished = {\url{https://huggingface.co/largitdata-inc/gemma-4-26b-a4b-it-fp8-dynamic-norouter}},
note = {Derived offline FP8 checkpoint from google/gemma-4-26B-A4B-it for vLLM serving, published by Largitdata Inc. \url{https://www.largitdata.com/}}
}
@misc{google_gemma4_26b_a4b_it,
title = {Gemma 4 26B-A4B IT},
author = {Google},
year = {2026},
howpublished = {\url{https://huggingface.co/google/gemma-4-26B-A4B-it}}
}
Disclaimer
Users are responsible for verifying license compatibility, downstream serving behavior, numerical quality, and safety characteristics for their own environment.
ไธญๆ่ชชๆ
็ตฆ vLLM ็จ็้ข็ท FP8 checkpoint โ ๆฏ BF16 ็ 47% VRAM๏ผๅนณ่ก่็่ฝๅๅค 80%ใ
ๆๅๅจ็ถฒ่ทฏไธๆพไบไธ่ผช๏ผๆฒๆๆพๅฐๅ ช็จ็ Gemma 4 26B ้ข็ท FP8 ็ๆฌ๏ผ็ดขๆง่ชๅทฑ vibe coding ๅไบไธ็๏ผ่ฒข็ป็ตฆ็คพ็พคใ
้ๅ Repo ๆไพๅพ google/gemma-4-26B-A4B-it ่ก็ๅบ็้ข็ท FP8 checkpoint๏ผ่ฎ vLLM ๅฏไปฅ็ดๆฅ่ผๅ
ฅๆๅ๏ผไธ้่ฆๅจๅๅๆๅท่ก on-the-fly ้ๅใ
็ฑ Largitdata Inc. ็ผไฝใ
ๆณจๆ๏ผ ้ๆฏ่ก็็ๆไฝ็จ checkpoint๏ผไธฆ้ Google ๅฎๆน็ผไฝใๅๅงๆจกๅ็ๆๆฌๆขๆฌพใๅฎๅ จๆๅผ่ๆไปถไปไปฅๅฎๆน็บๆบใ
ๆจกๅ็ดฐ็ฏ
- ๅบๅบๆจกๅ๏ผ
google/gemma-4-26B-A4B-it - ๆ ผๅผ๏ผ ้ข็ท FP8 checkpoint๏ผไพ vLLM ไฝฟ็จ
- ้ๅๅทฅๅ
ท๏ผ
llmcompressor - ้ๅๆนๅผ๏ผ
FP8_DYNAMIC - ๆ กๆบ่ณๆ๏ผ ไธ้่ฆ๏ผๅๆ ้ๅ๏ผ
- ๆ้ค็ๆฌ้๏ผ
norm้กไธ็ถญ tensor โ ้ฟๅ ้ๅ้ฉ่ญๆ็ข็expected 2D linear weight้ก้ฏ่ชคre:.*router\.proj$โ MoE router ๆฌ้๏ผ็ถญๆ่ Gemma4 vLLM ่ผๅ ฅ่ทฏๅพ็็ธๅฎนๆง
- ไธป่ฆ้จ็ฝฒ็ฎๆจ๏ผ
vllm/vllm-openai:gemma4
ๆธฌ่ฉฆ็ฐๅข
- GPU๏ผ
NVIDIA H200 NVL๏ผ143 GB VRAM๏ผ - Runtime๏ผ
vllm/vllm-openai:gemma4 - KV cache dtype๏ผ
fp8 max_model_len๏ผ32768gpu_memory_utilization๏ผ0.55
ๅๅๅฏฆๆธฌๆธๆ๏ผ
- ๆจกๅๆฌ้่ผๅ
ฅ๏ผ
15.76 s - ๆจกๅ่ผๅ
ฅ็ธฝ่จ๏ผ
16.88 s torch.compile๏ผ56.98 s- ๅผๆๅๅงๅ๏ผ
102.17 s /v1/modelsๅฐฑ็ท็ธฝๆ้๏ผ็ด153 s
ๅท่กๆๅฎน้๏ผ
max_num_batched_tokens = 8192- ๅฏ็จ KV cache ่จๆถ้ซ๏ผ
46.37 GiB - GPU KV cache ๅคงๅฐ๏ผ
405,184 tokens - ๆๅคงๅนณ่ก่็้๏ผ
32,768tokens/request๏ผ๏ผ38.87x
ๆๅๅฎน้ๆฏ่ผ
| ๆๆจ | FP8 Dynamic Norouter | BF16 ๅ็ |
|---|---|---|
| ๆจกๅ่ผๅ ฅ่จๆถ้ซ | 25.75 GiB | 48.5 GiB |
| GPU KV cache ๅคงๅฐ | 405,184 tokens | 225,376 tokens |
| ๆๅคงๅนณ่ก่็้ @ 32K tokens/req | 38.87x | 21.62x |
| VRAM ็ฏ็ | 47% | โ |
| KV cache ๅขๅ | 80% | โ |
ๅบ็คๆ่ฝๆธฌ่ฉฆ
ๅฎ่ซๆฑๆๆฉๆธฌ่ฉฆ๏ผOpenAI ็ธๅฎน vLLM endpoint๏ผ๏ผ
- prompt tokens๏ผ
38 - completion tokens๏ผ
256 - temperature๏ผ
0
| ๆๆจ | FP8 Dynamic Norouter | BF16 ๅ็ |
|---|---|---|
| ๅนณๅ็ซฏๅฐ็ซฏๅปถ้ฒ | 1.629 s | 1.536 s |
| ๅนณๅ completion ๅๅ้ | 157.19 tok/s | 166.62 tok/s |
| ๅนณๅ็ธฝๅๅ้ | 180.53 tok/s | 191.36 tok/s |
ไปฅไธ็บๅฎ่ซๆฑๆๆฉ่ทฏๅพๆธฌ้ๅผ๏ผ้ๅค็จๆถๅๅ้ๆธฌ่ฉฆใๅจ็็ข็ฐๅขๅค็จๆถๅ ดๆฏไธ๏ผFP8 ็ๆฌๆดๅคง็ KV cache ้ ๆ่ฝๆไพๆดๅฅฝ็ๆด้ซๅๅ้ใ
็ต่ซ๏ผBF16 ๅฎ่ซๆฑ็ฅๅฟซ๏ผ็ด 6%๏ผ๏ผไฝ FP8 ็ๆฌ VRAM ็จ้ๆธๅฐ 47%๏ผๅฏ็จ KV cache ๅขๅ 80%ใ้่ฆๅๆๆๅๅค็จๆถ็็็ข็ฐๅข๏ผFP8 ็ๆฌๆดๅ
ทๅชๅขใ
็ฒพๅบฆ่ฉไผฐ โ MMLU-Pro
ไฝฟ็จ็ธๅ harness ่ prompt template๏ผๅฐ FP8 checkpoint ่ BF16 ๅ็ๅๆๅท่ก MMLU-Pro๏ผๅ จ้๏ผ14 ้กๅฅ๏ผ๏ผ
| ๆจกๅ | MMLU-Pro ๆบ็ขบ็ | ็ธฝๅท่กๆ้ |
|---|---|---|
| 26B BF16๏ผๅบๆบ๏ผ | 81.59% | 65 min |
| 26B FP8 Dynamic Norouter๏ผๆฌ็ๆฌ๏ผ | 81.33% | 63 min |
| ฮ (FP8 โ BF16) | โ0.26 pp | โ |
FP8 ้ๅ็ไปฃๅนๅบๆฌ็บ้ถ โ 0.26 pp ็ๅทฎ่ท่ฝๅจ benchmark ๅช้ณ็ฏๅๅ งใๆๅ็ 26B BF16 ๆ็ธพ๏ผ81.59%๏ผๆฏ Google ๅฎๆน็ 82.6% ไฝ็ด 1.0 pp๏ผ้ๆฏ harness / prompt ๅทฎ็ฐ้ ๆ๏ผ้้ๅๅ้ก๏ผBF16 ่ FP8 ๅ ฉ่ ไธ่ตทๅ็งปใ
ไฝ็บๅฐ็ ง๏ผๅๆขไปถไธ 31B ็ๆฌ็ตๆ๏ผ
| ๆจกๅ | MMLU-Pro |
|---|---|
| 26B BF16 | 81.59% |
| 26B FP8 | 81.33% |
| 31B BF16 | 84.60% |
| 31B FP8 Dynamic | 84.55% |
ๅฆ่ฆ๏ผLargitData/gemma-4-31b-it-fp8ใ
ไฝฟ็จๆนๅผ
vLLM ๅๅ็ฏไพ๏ผ
docker run -d \
--name vllm-gemma4-26b-fp8-norouter \
--restart unless-stopped \
--ipc=host \
--shm-size 16G \
--gpus all \
-v /models \
-p 8001:8000 \
-e NVIDIA_VISIBLE_DEVICES=0 \
vllm/vllm-openai:gemma4 \
--model /models/gemma-4-26B-A4B-it-FP8-DYNAMIC-NOROUTER \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.55 \
--max-model-len 32768 \
--enable-auto-tool-choice \
--tool-call-parser gemma4 \
--host 0.0.0.0 \
--port 8000
ๅทฒ็ฅ้ๅถ
- ๅฎ่ซๆฑๅปถ้ฒๆฏ BF16 ้ซ็ด 6%๏ผไธปๅ ็บ FP8 dequantization ็้กๅค้้ท
- MMLU-Pro ๅทฒๅฎๆ๏ผ่ฆไธ๏ผ๏ผMT-Bench ็ญๅ ถไป benchmark ๅฐๆชๅท่ก๏ผๆญก่ฟ็คพ็พค่ฃๅ
- ๅ
ๅจ H200 NVL ไธๅฏฆๆธฌ๏ผๅ
ถไป GPU๏ผๅฆ A100ใH100๏ผๅฏ่ฝ้่ฆ่ชฟๆด
gpu-memory-utilization - MoE router ๆฌ้๏ผ
router.proj๏ผ่norm้กไธ็ถญ tensor ่ขซๆ้คๅจ้ๅ็ฏๅๅคไปฅ็ถญๆ vLLM ็ธๅฎนๆง๏ผ็ฎๅๆช่งๅฏๅฐๅๆตๅ่ณชไธ้๏ผไฝๅฐ็ก็ณป็ตฑๆง่ฉไผฐ
ไฝฟ็จๅ ดๆฏ
ๆญค checkpoint ้ฉ็จๆผ๏ผ
- ๅจ H200 ็ญ็ด็กฌ้ซไธไปฅ vLLM ้ฒ่ก็็ข้จ็ฝฒ
- ๅฏ้็พ็้ข็ท FP8 ๆๅๅฏฆ้ฉ
- ไธๅธๆๅจๅๅๆๅท่ก on-the-fly ้ๅ็็ฐๅข
- ้่ฆๆด้ซๅนณ่ก่็่ฝๅ็็็ขๆจ่ซๅ ดๆฏ
ๆญค checkpoint ไธๅไปฃๅๅงๅบๅบๆจกๅ็ๆไปถใๅฎๅ จๆๅผๆๆๆฌๆขๆฌพใ
ๆๆฌ
ๆญค Repo ๅ
ๅซๅบๆผ google/gemma-4-26B-A4B-it ็่ก็ checkpoint๏ผไฝฟ็จ้ ้ตๅฎ Gemma ไฝฟ็จๆขๆฌพใ
ๅ ่ฒฌ่ฒๆ
ไฝฟ็จ่ ้่ช่ก้ฉ่ญๆๆฌ็ธๅฎนๆงใไธๆธธๆๅ่ก็บใๆธๅผๅ่ณช่ๅฎๅ จ็นๆงใ
- Downloads last month
- 34,254