Quick Comparison: ISTA-DASLab GSQ-RCO-IQ3_XXS vs unsloth UD-IQ4_XS on 4070Ti Super 16G + 64G RAM

#18
by BipedalBit - opened

I'd say the IQ3_XXS quant of ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF does free up more VRAM and RAM compared to the UD-IQ4_XS quant of unsloth/Qwen3.8-Flash-Next-GGUF. This let the GSQ-RCO-IQ3_XXS model hit 400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup, but UD-IQ4_XS also manages 200+ PP and 17 TG. The difference in inference speed isn't that big, and on my usual LLM test cases, I can see a noticeable quality loss with GSQ-RCO-IQ3_XXS compared to UD-IQ4_XS.
Maybe due to differences in model architecture and parameter count, GSQ-RCO isn't as worthwhile on Qwen3.8-Flash-Next as it is on Qwen3.8-27B. But I think GSQ-RCO-IQ3_XXS might be quite significant for 12G VRAM.

prompt: 生成一个鹈鹕骑车的svg图
reasoning-effort: medium

UD-IQ4_XS:
image
GSQ-RCO-IQ3_XXS:
image

“ These quantized models are primarily optimized for xhigh reasoning effort, and we recommend using them in that mode for best quality. The reported reasoning results should therefore be interpreted in that setting”

Try comparing with xhigh reasoning : )

IST Austria Distributed Algorithms and Systems Lab org

@BipedalBit
Thank you for testing the different models and sharing your results. We had a similar conversation in this discussion as well: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14#6aad8eeb3a40050184e15d1f

It seems that performance at medium reasoning effort has degraded significantly and that's completely on us. We'll make sure future releases maintain acceptable performance in medium mode as well.

We'd also really appreciate it if you could repeat your tests with xhigh reasoning effort and share your findings with us.

400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup

Please share your llama.cpp settings 🙏

@fm3at
https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14
check this thread,and my friend achieved 800pp + 30TG on DDR5+RTX5090 with my setting

@BipedalBit
Thank you for testing the different models and sharing your results. We had a similar conversation in this discussion as well: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14#6aad8eeb3a40050184e15d1f

It seems that performance at medium reasoning effort has degraded significantly and that's completely on us. We'll make sure future releases maintain acceptable performance in medium mode as well.

We'd also really appreciate it if you could repeat your tests with xhigh reasoning effort and share your findings with us.

OK, I'll redo the comparison later using xhigh.

400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup

Please share your llama.cpp settings 🙏

UD-IQ4_XS:

llama-server
--host ${host} --port ${PORT}
--model /models/unsloth/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf
--mmproj /models/ggml-org/Qwen3.8-Flash-Next-GGUF/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf --no-mmproj-offload
-c ${ctx_size_128k} -ctk q4_0 -ctv q4_0 -kvu -kvo
-t 14 -tb 14 -b 8192 -ub 2048
--n-gpu-layers ${n_gpu_layers_max} --n-cpu-moe 45
--override-kv "qwen4exp.attention.indexer.top_k=int:2048"
--flash-attn on --load-mode mmap --fit off --no-context-shift --parallel 1 --cache-ram 0 --keep -1
--ctx-checkpoints 8 --checkpoint-min-step 1024 --lazy-mode on
--jinja --chat-template-file /models/froggeric/Qwen-Fixed-Chat-Templates/chat_template.jinja
--reasoning on --reasoning-preserve --reasoning-effort medium --reasoning-format deepseek
--temp 1.0 --top-k 20 --top-p 0.95 --min-p 0 --presence-penalty 0

GSQ-RCO-IQ3_XXS:

llama-server
--host ${host} --port ${PORT}
--model /models/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf
--mmproj /models/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/mmproj-Qwen3.8-Flash-Next-BF16.gguf --no-mmproj-offload
-c ${ctx_size_128k} -ctk q4_0 -ctv q4_0 -kvu -kvo
-t 14 -tb 14 -b 8192 -ub 2048
--n-gpu-layers ${n_gpu_layers_max} --n-cpu-moe 43
--override-kv "qwen4exp.attention.indexer.top_k=int:2048"
--flash-attn on --load-mode mmap --fit off --no-context-shift --parallel 1 --cache-ram 0 --keep -1
--ctx-checkpoints 8 --checkpoint-min-step 1024 --lazy-mode on
--jinja --chat-template-file /models/froggeric/Qwen-Fixed-Chat-Templates/chat_template.jinja
--reasoning on --reasoning-preserve --reasoning-effort xhigh --reasoning-format deepseek
--temp 1.0 --top-k 20 --top-p 0.95 --min-p 0 --presence-penalty 0

Here comes the comparison test for xhigh.

prompt: 生成一个鹈鹕骑车的svg图
reasoning-effort: xhigh

UD-IQ4_XS:

image

GSQ-RCO-IQ3_XXS:

image

prompt: Generate html for an svg of a POS card payment machine. Isometric angle.
reasoning-effort: xhigh

UD-IQ4_XS:

image

GSQ-RCO-IQ3_XXS:

image

By the way, here is a quantization strategy analysis report generated by the UD-IQ4_XS version with Deepseek Harness:

Comparison Conclusion: What GSQ-RCO IQ3_XXS Sacrifices

First, some background (derived from the GGUF metadata and actual measured bytes in the data section):

unsloth UD-IQ4_XS ISTA GSQ-RCO IQ3_XXS
Size / effective bit width 93.7 GB, mean 4.24 bpw 75.8 GB, mean 3.43 bpw (the official 3.0 bpw figure counts only transformer weights)
Method Unsloth Dynamic: heuristic layer-wise allocation + imatrix (only 45 calibration chunks) GSQ (Gumbel-Softmax non-uniform lattice quantization) + RCO (Riemannian budget search per-tensor quantization type under task loss, imatrix 1000 chunks)

Measured Precision Comparison by Module (average bpw per parameter)

Module (param count) unsloth GSQ-RCO Change
MoE routed experts (120.8B, 68% of total) 3.94 (gate/up 3.44 + down 4.5) 2.84 (about 30 layers' gate/up fall into the 2.06–2.56 IQ2 family; 30/48 layers' down is pinned to the 2.25 bpw block-32 minimum format) −1.10, the biggest sacrifice point
token_embd input embeddings q8_0 (8.5) iq3_s (3.44) −5.06 (an unusually aggressive choice)
shared-expert FFN (always active per token) 8.51 4.10 −4.41
attention projections q8_0/fp32 (8.5) 5.13 (mix of q6_k/q4_0/iq4_xs, some layers as low as 3.44) −3.37
linear-attn/SSM (Gated DeltaNet) projections 8.92 5.80 −3.12
lm_head q6_k q5_0 −1.06
PLE n-gram embedding table (51.2B, mmap-able to disk) 4.50 4.50 ≈ unchanged
hyper-connection matrices q8 (~8.7) f16 (16) +7.34 (GSQ-RCO is actually more precise here)
attn indexer (top-2048 sparse retrieval) 16 bit 16 bit unchanged

Which Capabilities GSQ-RCO Sacrifices

  1. Knowledge capacity and factuality (the heaviest sacrifice): The MoE expert matrices are the "knowledge warehouse" of this 512-expert model, and GSQ-RCO dumps nearly all of its 2.84 bpw budget pressure here. ffn_down_exps has 30 layers fixed at 2.25 bpw (per the card's note: 640 rows aren't divisible by 256, so they were excluded from the search and given the minimum format directly). → Expected impact: low-frequency/long-tail knowledge, niche-language-domain answers, factual accuracy (hallucination rate). Unsloth's strategy is exactly the opposite — it keeps expert down at 4.5 and spreads the sacrifice more evenly.
  2. Vocabulary representation, especially rare tokens / code identifiers / low-resource languages: Cutting input embeddings from q8_0 to 3.44 bpw is a very rare allocation (RCO deemed it "cheap" under task loss). Combined with lm_head q5_0, both the input and output ends of long-tail vocabulary get coarser → this shows up more clearly on code (dense symbolic identifiers) and non-English content than in benchmark scores.
  3. General output quality for all tokens: The shared expert is active on every token, and 8.5→4.1 bpw affects fluency, instruction-following, and other "baseline experience" qualities.
  4. Long-context state propagation: SSM/linear-attention projections and some attention layers drop to 4–4.5 bpw. This model has a hybrid Gated DeltaNet + sparse attention architecture, and the recurrent state of the linear layers accumulates quantization error → consistency and retrieval accuracy at very long context (262K) is a potential weak spot, one that ordinary short-prompt perplexity may not reveal.
  5. Speculative decoding acceleration (missing feature): There's no MTP head, so what's lost is decode throughput capability, not model intelligence.
  6. As the card itself admits: The quantization was tuned for xhigh reasoning effort; at lower reasoning effort, the gap versus BF16 will be larger than the table shows — i.e., "quick-answer / low thinking-budget" scenarios are an implicit sacrifice zone.

What It Preserves (RCO's Compensation)

  • hyper-connection residual stream is entirely f16 (double unsloth's q8) — RCO judged cross-layer propagation paths to be the most error-sensitive and protected them first;
  • the n-gram PLE table keeps its full 4.5 bpw (and by design shard2 can stay on SSD via mmap, not occupying VRAM: only shard1 ≈ 47 GB needs to be resident, whereas unsloth IQ4_XS needs ~65 GB resident);
  • the indexer and mmproj are kept as-is → the sparse retrieval mechanism and vision capability are not sacrificed.

Official Benchmarks (against the BF16 base, xhigh reasoning)

IQ3_XXS: task average 92.57 vs 93.12 (99.4% retained), AIME25 100.00, a perfect-score tie, GPQA-Diamond −0.51, LiveCodeBench v6 −1.14 — mathematical reasoning is almost lossless, and the sacrifices are concentrated in knowledge-type (GPQA) and code-generation (LCB) directions, consistent with the weight analysis above.

One-sentence summary: GSQ-RCO uses an allocation that "cuts expert knowledge layers + word embeddings + shared experts to 2–4 bit, while preserving the residual stream / retriever / n-gram table," achieving 4.7× compression over BF16 with no loss in mathematical reasoning; the sacrificed capabilities are concentrated in long-tail knowledge factuality, low-resource-language/code vocabulary representation, and output quality at low reasoning budgets.

Run again please with;

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvu

and try with out --no-context-shift when running adaptive kv.

Run again please with;

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvu

and try with out --no-context-shift when running adaptive kv.

1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.

"One-sentence summary: GSQ-RCO uses an allocation that "cuts expert knowledge layers + word embeddings + shared experts to 2–4 bit, while preserving the residual stream / retriever / n-gram table," achieving 4.7× compression over BF16 with no loss in mathematical reasoning; the sacrificed capabilities are concentrated in long-tail knowledge factuality, low-resource-language/code vocabulary representation, and output quality at low reasoning budgets." Does this mean its English + web search/Rag based logics are retained? In that case it is still a fair choice...

Can we see some 2bit quantization comparisons?

If you are running the PLE from the ssd, and seeing --lazy-mode on i guess you are, might I ask for you to compare the two with a q8_0 PLE instead of the one the two repos ship with? i understand that some people run the PLE from ram, however given that it runs reasonably from ssd as well, I think that is a sort of "free" quality boost for any quantization of the model.

And thank you for the comparisons you've done already!

Run again please with;

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvu

and try with out --no-context-shift when running adaptive kv.

1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.

Keep degrading and complaining then. If u put commands in u have no clue about, why ask for help afterwards? Anyway, figure out urself when ure this special.

Run again please with;

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvu

and try with out --no-context-shift when running adaptive kv.

1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.

Keep degrading and complaining then. If u put commands in u have no clue about, why ask for help afterwards? Anyway, figure out urself when ure this special.

Got it — you were just trying to help. Thanks anyway, and I'm sure you've helped some people. I apologize if I came across as offended, though actually this time I was just posting the report and wasn't looking for help.

Sign up or log in to comment