Instructions to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS # Run inference directly in the terminal: ./llama-cli -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
- LM Studio
- Jan
- vLLM
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
- Ollama
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
- Unsloth Desktop
- Pi
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
- Lemonade
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-GSQ-RCO-GGUF-IQ2_XS
List all available models
lemonade list
- Hermes Agent
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ2_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Comparison: ISTA-DASLab GSQ-RCO-IQ3_XXS vs unsloth UD-IQ4_XS on 4070Ti Super 16G + 64G RAM
I'd say the IQ3_XXS quant of ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF does free up more VRAM and RAM compared to the UD-IQ4_XS quant of unsloth/Qwen3.8-Flash-Next-GGUF. This let the GSQ-RCO-IQ3_XXS model hit 400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup, but UD-IQ4_XS also manages 200+ PP and 17 TG. The difference in inference speed isn't that big, and on my usual LLM test cases, I can see a noticeable quality loss with GSQ-RCO-IQ3_XXS compared to UD-IQ4_XS.
Maybe due to differences in model architecture and parameter count, GSQ-RCO isn't as worthwhile on Qwen3.8-Flash-Next as it is on Qwen3.8-27B. But I think GSQ-RCO-IQ3_XXS might be quite significant for 12G VRAM.
prompt: 生成一个鹈鹕骑车的svg图
reasoning-effort: medium
“ These quantized models are primarily optimized for xhigh reasoning effort, and we recommend using them in that mode for best quality. The reported reasoning results should therefore be interpreted in that setting”
Try comparing with xhigh reasoning : )
@BipedalBit
Thank you for testing the different models and sharing your results. We had a similar conversation in this discussion as well: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14#6aad8eeb3a40050184e15d1f
It seems that performance at medium reasoning effort has degraded significantly and that's completely on us. We'll make sure future releases maintain acceptable performance in medium mode as well.
We'd also really appreciate it if you could repeat your tests with xhigh reasoning effort and share your findings with us.
400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup
Please share your llama.cpp settings 🙏
@fm3at
https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14
check this thread,and my friend achieved 800pp + 30TG on DDR5+RTX5090 with my setting
@BipedalBit
Thank you for testing the different models and sharing your results. We had a similar conversation in this discussion as well: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/14#6aad8eeb3a40050184e15d1fIt seems that performance at medium reasoning effort has degraded significantly and that's completely on us. We'll make sure future releases maintain acceptable performance in medium mode as well.
We'd also really appreciate it if you could repeat your tests with xhigh reasoning effort and share your findings with us.
OK, I'll redo the comparison later using xhigh.
400+ PP and 19 TG on my Nvidia 4070Ti Super 16G VRAM + 64G RAM + NVMe SSD setup
Please share your llama.cpp settings 🙏
UD-IQ4_XS:
llama-server
--host ${host} --port ${PORT}
--model /models/unsloth/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf
--mmproj /models/ggml-org/Qwen3.8-Flash-Next-GGUF/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf --no-mmproj-offload
-c ${ctx_size_128k} -ctk q4_0 -ctv q4_0 -kvu -kvo
-t 14 -tb 14 -b 8192 -ub 2048
--n-gpu-layers ${n_gpu_layers_max} --n-cpu-moe 45
--override-kv "qwen4exp.attention.indexer.top_k=int:2048"
--flash-attn on --load-mode mmap --fit off --no-context-shift --parallel 1 --cache-ram 0 --keep -1
--ctx-checkpoints 8 --checkpoint-min-step 1024 --lazy-mode on
--jinja --chat-template-file /models/froggeric/Qwen-Fixed-Chat-Templates/chat_template.jinja
--reasoning on --reasoning-preserve --reasoning-effort medium --reasoning-format deepseek
--temp 1.0 --top-k 20 --top-p 0.95 --min-p 0 --presence-penalty 0
GSQ-RCO-IQ3_XXS:
llama-server
--host ${host} --port ${PORT}
--model /models/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf
--mmproj /models/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/mmproj-Qwen3.8-Flash-Next-BF16.gguf --no-mmproj-offload
-c ${ctx_size_128k} -ctk q4_0 -ctv q4_0 -kvu -kvo
-t 14 -tb 14 -b 8192 -ub 2048
--n-gpu-layers ${n_gpu_layers_max} --n-cpu-moe 43
--override-kv "qwen4exp.attention.indexer.top_k=int:2048"
--flash-attn on --load-mode mmap --fit off --no-context-shift --parallel 1 --cache-ram 0 --keep -1
--ctx-checkpoints 8 --checkpoint-min-step 1024 --lazy-mode on
--jinja --chat-template-file /models/froggeric/Qwen-Fixed-Chat-Templates/chat_template.jinja
--reasoning on --reasoning-preserve --reasoning-effort xhigh --reasoning-format deepseek
--temp 1.0 --top-k 20 --top-p 0.95 --min-p 0 --presence-penalty 0
By the way, here is a quantization strategy analysis report generated by the UD-IQ4_XS version with Deepseek Harness:
Comparison Conclusion: What GSQ-RCO IQ3_XXS Sacrifices
First, some background (derived from the GGUF metadata and actual measured bytes in the data section):
unsloth UD-IQ4_XS |
ISTA GSQ-RCO IQ3_XXS |
|
|---|---|---|
| Size / effective bit width | 93.7 GB, mean 4.24 bpw | 75.8 GB, mean 3.43 bpw (the official 3.0 bpw figure counts only transformer weights) |
| Method | Unsloth Dynamic: heuristic layer-wise allocation + imatrix (only 45 calibration chunks) | GSQ (Gumbel-Softmax non-uniform lattice quantization) + RCO (Riemannian budget search per-tensor quantization type under task loss, imatrix 1000 chunks) |
Measured Precision Comparison by Module (average bpw per parameter)
| Module (param count) | unsloth | GSQ-RCO | Change |
|---|---|---|---|
| MoE routed experts (120.8B, 68% of total) | 3.94 (gate/up 3.44 + down 4.5) | 2.84 (about 30 layers' gate/up fall into the 2.06–2.56 IQ2 family; 30/48 layers' down is pinned to the 2.25 bpw block-32 minimum format) | −1.10, the biggest sacrifice point |
| token_embd input embeddings | q8_0 (8.5) | iq3_s (3.44) | −5.06 (an unusually aggressive choice) |
| shared-expert FFN (always active per token) | 8.51 | 4.10 | −4.41 |
| attention projections | q8_0/fp32 (8.5) | 5.13 (mix of q6_k/q4_0/iq4_xs, some layers as low as 3.44) | −3.37 |
| linear-attn/SSM (Gated DeltaNet) projections | 8.92 | 5.80 | −3.12 |
| lm_head | q6_k | q5_0 | −1.06 |
| PLE n-gram embedding table (51.2B, mmap-able to disk) | 4.50 | 4.50 | ≈ unchanged |
| hyper-connection matrices | q8 (~8.7) | f16 (16) | +7.34 (GSQ-RCO is actually more precise here) |
| attn indexer (top-2048 sparse retrieval) | 16 bit | 16 bit | unchanged |
Which Capabilities GSQ-RCO Sacrifices
- Knowledge capacity and factuality (the heaviest sacrifice): The MoE expert matrices are the "knowledge warehouse" of this 512-expert model, and GSQ-RCO dumps nearly all of its 2.84 bpw budget pressure here.
ffn_down_expshas 30 layers fixed at 2.25 bpw (per the card's note: 640 rows aren't divisible by 256, so they were excluded from the search and given the minimum format directly). → Expected impact: low-frequency/long-tail knowledge, niche-language-domain answers, factual accuracy (hallucination rate). Unsloth's strategy is exactly the opposite — it keeps expert down at 4.5 and spreads the sacrifice more evenly. - Vocabulary representation, especially rare tokens / code identifiers / low-resource languages: Cutting input embeddings from q8_0 to 3.44 bpw is a very rare allocation (RCO deemed it "cheap" under task loss). Combined with lm_head q5_0, both the input and output ends of long-tail vocabulary get coarser → this shows up more clearly on code (dense symbolic identifiers) and non-English content than in benchmark scores.
- General output quality for all tokens: The shared expert is active on every token, and 8.5→4.1 bpw affects fluency, instruction-following, and other "baseline experience" qualities.
- Long-context state propagation: SSM/linear-attention projections and some attention layers drop to 4–4.5 bpw. This model has a hybrid Gated DeltaNet + sparse attention architecture, and the recurrent state of the linear layers accumulates quantization error → consistency and retrieval accuracy at very long context (262K) is a potential weak spot, one that ordinary short-prompt perplexity may not reveal.
- Speculative decoding acceleration (missing feature): There's no MTP head, so what's lost is decode throughput capability, not model intelligence.
- As the card itself admits: The quantization was tuned for xhigh reasoning effort; at lower reasoning effort, the gap versus BF16 will be larger than the table shows — i.e., "quick-answer / low thinking-budget" scenarios are an implicit sacrifice zone.
What It Preserves (RCO's Compensation)
- hyper-connection residual stream is entirely f16 (double unsloth's q8) — RCO judged cross-layer propagation paths to be the most error-sensitive and protected them first;
- the n-gram PLE table keeps its full 4.5 bpw (and by design shard2 can stay on SSD via mmap, not occupying VRAM: only shard1 ≈ 47 GB needs to be resident, whereas unsloth IQ4_XS needs ~65 GB resident);
- the indexer and mmproj are kept as-is → the sparse retrieval mechanism and vision capability are not sacrificed.
Official Benchmarks (against the BF16 base, xhigh reasoning)
IQ3_XXS: task average 92.57 vs 93.12 (99.4% retained), AIME25 100.00, a perfect-score tie, GPQA-Diamond −0.51, LiveCodeBench v6 −1.14 — mathematical reasoning is almost lossless, and the sacrifices are concentrated in knowledge-type (GPQA) and code-generation (LCB) directions, consistent with the weight analysis above.
One-sentence summary: GSQ-RCO uses an allocation that "cuts expert knowledge layers + word embeddings + shared experts to 2–4 bit, while preserving the residual stream / retriever / n-gram table," achieving 4.7× compression over BF16 with no loss in mathematical reasoning; the sacrificed capabilities are concentrated in long-tail knowledge factuality, low-resource-language/code vocabulary representation, and output quality at low reasoning budgets.
Run again please with;
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvu
and try with out --no-context-shift when running adaptive kv.
Run again please with;
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvuand try with out --no-context-shift when running adaptive kv.
1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.
"One-sentence summary: GSQ-RCO uses an allocation that "cuts expert knowledge layers + word embeddings + shared experts to 2–4 bit, while preserving the residual stream / retriever / n-gram table," achieving 4.7× compression over BF16 with no loss in mathematical reasoning; the sacrificed capabilities are concentrated in long-tail knowledge factuality, low-resource-language/code vocabulary representation, and output quality at low reasoning budgets." Does this mean its English + web search/Rag based logics are retained? In that case it is still a fair choice...
Can we see some 2bit quantization comparisons?
If you are running the PLE from the ssd, and seeing --lazy-mode on i guess you are, might I ask for you to compare the two with a q8_0 PLE instead of the one the two repos ship with? i understand that some people run the PLE from ram, however given that it runs reasonably from ssd as well, I think that is a sort of "free" quality boost for any quantization of the model.
And thank you for the comparisons you've done already!
Run again please with;
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvuand try with out --no-context-shift when running adaptive kv.
1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.
Keep degrading and complaining then. If u put commands in u have no clue about, why ask for help afterwards? Anyway, figure out urself when ure this special.
Run again please with;
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
ctk q8_0 -ctv q4_0 -kvuand try with out --no-context-shift when running adaptive kv.
1、For GGML_CUDA_ENABLE_UNIFIED_MEMORY, hey bro, my setup doesn't have unified memory, and that has nothing to do with model quality anyway.
2、For ctk q8_0, in my tests, both unsloth and GSQ-RCO used ctk q4_0, which I think is fair. If you just want to see the effect of a specific configuration, I guess you'd probably need to try it yourself.
3、For --no-context-shift,I can confirm that my test cases did not cause the context to fill up. Also, the default value of --no-context-shift is disabled, so removing this option won't change anything.Keep degrading and complaining then. If u put commands in u have no clue about, why ask for help afterwards? Anyway, figure out urself when ure this special.
Got it — you were just trying to help. Thanks anyway, and I'm sure you've helped some people. I apologize if I came across as offended, though actually this time I was just posting the report and wasn't looking for help.





