will you make a q8?

#1
by ahinks - opened

will you make a q8 quant i prefer running q8 for quality?
thank you!

I'll see what I can do

will you make a q8 quant i prefer running q8 for quality?
thank you!

It's up!

Its definitely NOT as good IMO something is happening.... the scores are BS imo thsi was q8 and q8 on a strix halo
Untitled

Mine is a true ROCmFPX Q8 across the board, every tensor including the output layer. ROCmFPX's 8-bit isn't a drop-in equivalent of standard Q8_0. It's a different format with its own rounding, so the two should land close but not identical. Some difference between them is expected rather than a sign that something's wrong.

Depending on what you're comparing against, it may also not be a fair comparison. DavidAU's quants keep the output tensor at full precision, so his are doing something different at the same nominal bit level.

I'm regenerating a Q8 now that retains more quality. It'll be a larger file, more on track with his GGUFs. I'll swap it in when it's done.

Thank you for the insight

we are all just vibe benching too!

the prompt i used is the typical 'generate a SVG of a pelican riding a bicycle"

I ran with thinking off multiple seeds and thinking on multiple seeds no token limit.
This Quant does have a speedup compared! I typically run Q8 when it makes sence strix halo feature dedicated hardware instructions optimized for INT8 (8-bit integer) operations. Processing Q8 weights cuts total memory bandwidth, still slower then Q4 but most of the time running q8 is just better on my hardware.

thank you for the effort! and the resources!

Thank you for the insight

we are all just vibe benching too!

the prompt i used is the typical 'generate a SVG of a pelican riding a bicycle"

I ran with thinking off multiple seeds and thinking on multiple seeds no token limit.
This Quant does have a speedup compared! I typically run Q8 when it makes sence strix halo feature dedicated hardware instructions optimized for INT8 (8-bit integer) operations. Processing Q8 weights cuts total memory bandwidth, still slower then Q4 but most of the time running q8 is just better on my hardware.

thank you for the effort! and the resources!

New version is up. https://huggingface.co/lmcoleman/Qwen3.6-27B-Fable-Fusion-711-MTP-ROCmFPX-GGUF/resolve/main/Qwen3.6-27B-Fable-Fusion-711-MTP-Q8_0_ROCMFPX-BF16Out-imatrix.gguf

It reduced the quantization loss by 63% and upped the disk space by 1.1GiB. Loss is now only 0.16% against BF16.

Significant improvement for a huge slowdown thank you !

pelican-bf16out-vs-q8

Sign up or log in to comment