Gemma 4 NVFP4

#10
by Zek-Takai - opened

Hello, For any that are interested I managed to make a nvfp4 quant of Gemma 26b a4b it with the heads in two quant states. The Attention layers are at nvfp4 and the global layers are at fp8. Current test runs on the dgx spark with vllm .26 and Hermes has the decode tokens holding at high inference even at 60k context without a mtp head. https://huggingface.co/Zek-Takai/gemma-4-26B-A4B-it-NVFP4-fp8global

dsikka changed discussion status to closed

Sign up or log in to comment