Instructions to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf cafonez/Qwen3.8-Flash-Next-HC-Q8 # Run inference directly in the terminal: llama cli -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf cafonez/Qwen3.8-Flash-Next-HC-Q8 # Run inference directly in the terminal: llama cli -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf cafonez/Qwen3.8-Flash-Next-HC-Q8 # Run inference directly in the terminal: ./llama-cli -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf cafonez/Qwen3.8-Flash-Next-HC-Q8 # Run inference directly in the terminal: ./build/bin/llama-cli -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Use Docker
docker model run hf.co/cafonez/Qwen3.8-Flash-Next-HC-Q8
- LM Studio
- Jan
- vLLM
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cafonez/Qwen3.8-Flash-Next-HC-Q8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cafonez/Qwen3.8-Flash-Next-HC-Q8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/cafonez/Qwen3.8-Flash-Next-HC-Q8
- Ollama
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with Ollama:
ollama run hf.co/cafonez/Qwen3.8-Flash-Next-HC-Q8
- Unsloth Desktop
- Pi
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "cafonez/Qwen3.8-Flash-Next-HC-Q8" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with Docker Model Runner:
docker model run hf.co/cafonez/Qwen3.8-Flash-Next-HC-Q8
- Lemonade
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull cafonez/Qwen3.8-Flash-Next-HC-Q8
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-HC-Q8-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default cafonez/Qwen3.8-Flash-Next-HC-Q8
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use cafonez/Qwen3.8-Flash-Next-HC-Q8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cafonez/Qwen3.8-Flash-Next-HC-Q8
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "cafonez/Qwen3.8-Flash-Next-HC-Q8" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8 Flash Next HC-Q8
Quantized GGUF of Qwen/Qwen3.8-Flash-Next. This is not an official Qwen release. The weights remain under the Qwen Community License 1.0.
The file is Qwen3.8-Flash-Next-HC-Q8.gguf.
- 179.551 billion parameters
- 92,347,778,272 bytes
- 4.11 bits per weight on average
- Architecture
qwen4exp, context trained at 262,144 - The MTP head is inside this file
| Weights | Parameters | Storage | Bits per weight |
|---|---|---|---|
| Network, including experts | 126.33B | Q4_0_ROCMI4 | 4.25 |
| N-gram table | 51.20B | Q3_0_ROCMFPX | 3.50 |
| Hyperconnection up and down | 0.655B | Q8_0 | 8.50 |
| Injection and other BF16 tensors | 0.661B | BF16 | 16 |
| One remaining tensor | 0.636B | Q6_K | 6.56 |
HC-Q8 starts from a Q4_0_ROCMI4 quant. The 200 hyperconnection up and down matrices were restored from the original BF16 tensors in Qwen revision de4b8e4d and stored as ordinary Q8_0. The 98 injection tensors stayed BF16. Every other tensor kept the ROCmI4 bytes.
Q4_0_ROCMI4 and Q3_0_ROCMFPX are ROCmFPX quant types. Stock llama.cpp does not load this file.
Run it with ROCmFPX HIP
Use a HIP llama-server built from ROCmFPX/ROCmFPX, with W4A4 disabled for this profile. Some ROCmFPX documentation names a serving preset or model generation "ROCmFP4 v2"; that is not a ROCmFPX runtime version requirement. Do not pass a separate MTP file: this GGUF already contains the MTP head. The separate-file FP4 package uses -md; that flag does not belong on this quant.
HSA_OVERRIDE_GFX_VERSION=11.5.1 is for Strix Halo (gfx1151, RDNA 3.5). Leave it unset on RDNA 3 and on RDNA 4 cards such as the Radeon AI Pro R9700.
Strix Halo HIP graph settings
On Strix Halo, use a ROCmFPX HIP build that includes the captured-copy support. To keep HIP graph capture enabled, set the opt-in copy-kernel switch and make sure graph disabling is unset:
export ROCMFPX_HIP_GRAPH_COPY_KERNEL=1
unset GGML_CUDA_DISABLE_GRAPHS
The copy-kernel switch replaces captured scalar device-to-device copies on gfx1151; it does not turn graph capture on by itself. This implementation is present in ROCmFPX main at commit 721db4193c9736c4d48f508993fcdfc6c75e180a and later builds that retain it.
If you still see repetitive or corrupted output with graphs enabled, use this correctness-first fallback:
unset ROCMFPX_HIP_GRAPH_COPY_KERNEL
export GGML_CUDA_DISABLE_GRAPHS=1
These settings are alternatives. GGML_CUDA_DISABLE_GRAPHS=1 disables the capture that activates the copy-kernel switch, so do not set it when testing the graph-enabled path. Disabling graphs may reduce decode speed. The reported failure was specific to a Strix Halo HIP runtime/model setup; verify output on your own build and configuration.
# gfx1151 / Strix Halo only
export HSA_OVERRIDE_GFX_VERSION=11.5.1
export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
export GGML_HIP_ENABLE_UNIFIED_MEMORY=1
export MALLOC_ARENA_MAX=2
llama-server \
-m Qwen3.8-Flash-Next-HC-Q8.gguf \
--alias qwen38-flash-next-hcq8 \
--host 127.0.0.1 --port 8128 \
--lazy-mode on-direct \
-dev ROCm0 -ngl 99 -ngld 99 \
-c 65536 -np 1 \
-b 2048 -ub 512 -t 16 -tb 16 \
-fa on -ctk f16 -ctv f16 \
--fit off --jinja --reasoning-format deepseek \
--no-kv-unified --no-context-shift --cont-batching \
--cache-prompt --cache-idle-slots --cache-reuse 0 --cache-ram 1024 \
--ctx-checkpoints 32 --checkpoint-min-step 8192 \
--metrics \
--spec-type ngram-mod,draft-mtp \
--spec-draft-device ROCm0 --spec-draft-ngl 99 \
--spec-draft-n-max 3 --spec-draft-p-min 0 --spec-draft-p-split 0.10 \
--no-spec-draft-backend-sampling \
--spec-ngram-mod-n-match 16 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 64
Context length
The command above follows the ROCmFPX v2 profile, so it reserves a 65,536-token KV cache. This GGUF is the 262,144-token model either way. A prompt that fits in 65,536 runs the same forward pass when the server is started at 65,536 or at 262,144.
llama-server sizes that cache from -c at startup. A server started at 65,536 stops there. To serve a longer prompt, change that flag and start the server again, up to the trained length:
-c 262144
Keep -ctk f16 -ctv f16. The F16 cache for all 262,144 tokens is most of the memory on a 128 GB machine. This file, served at -c 262144 on a 128 GB Strix Halo, used about 106 GiB and left about 10 GiB free. Use a shorter -c when the machine does not have that much memory.
The 32k, 64k, 128k, and 262k checks were different fills of one server already started at 262,144. Decode with 262,012 tokens in cache was 33 tokens/s.
On that server, one uninterrupted prefill of 262,016 tokens aborted once the key length passed 262,140 inside an 8192-token batch, with QSA decode enabled. Filling 253,824 tokens and then extending the cache in steps under 128 tokens completed, and decode ran. The v2 command on this card uses -b 2048 -ub 512.
Images need the original F16 vision projector from the Qwen repository beside this file, selected with --mmproj. That projector is not in this repo.
- Downloads last month
- 2,112
We're not able to determine the quantization variants.
Model tree for cafonez/Qwen3.8-Flash-Next-HC-Q8
Base model
Qwen/Qwen3.8-Flash-Next