Instructions to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0 # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0 # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0 # Run inference directly in the terminal: ./llama-cli -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
- LM Studio
- Jan
- vLLM
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "agentionai/Qwen3.8-Flash-Next-Gyro-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentionai/Qwen3.8-Flash-Next-Gyro-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
- Ollama
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with Ollama:
ollama run hf.co/agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
- Unsloth Desktop
- Pi
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with Docker Model Runner:
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
- Lemonade
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-Gyro-GGUF-TQ2_0
List all available models
lemonade list
- Hermes Agent
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use agentionai/Qwen3.8-Flash-Next-Gyro-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:TQ2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next Gyro โ GGUF (experimental)
Qwen3.8-Flash-Next on a single GPU: one model, three sizes.
| File | For | GPU memory (64k context) | Status |
|---|---|---|---|
Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf |
32 GB cards | 28.8 GiB | available |
Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf |
48 GB: 2ร24 GB cards, 64 GB unified-memory machines | ~37 GiB (34.35 GiB weights) | available: KLD 0.306 (S: 0.435), top-1 78.1% |
Qwen3.8-Flash-Next-Gyro-L.gguf |
64 GB cards and unified-memory machines | ~46โ50 GiB | planned |
All three keep the model's architecture; the larger files give the routed experts and the shared layers more bits.
The sizes for unified-memory machines (such as Strix Halo) leave room for the operating system. All three need our
llama.cpp build with Vulkan (see Quick start). The TQ1_0 / TQ2_0 in the file names is a size class for the
Hub's file browser: the files use our own rotor formats, not the standard ternary types, and stock llama.cpp cannot
load them.
The rest of this card describes Gyro-S: 28.8 GiB on the GPU at 64k context (our Q4: 67 GiB) | no loops in our looping test | with MTP drafting on one R9700 32 GB: 71 tok/s generating code, 86 tok/s editing supplied code (58 without drafting; 1,246 tok/s prefill). Its calibration set is weighted toward code and agentic work. This is an experimental release for people who want to try it and tell us how it behaves. The format and kernels may still change.
Highlights
- Fits a 32 GB card. The mixture-of-experts model (126B parameters, 121B of them in the routed experts) needs 28.8 GiB of GPU memory with 64k of context (see Memory and context). Its 51.2B-parameter n-gram table stays on disk and is read row by row as needed.
- Reasons without looping. Low-bit quants of this model tend to loop in long reasoning. In our test to induce looping (see Behaviour) this file looped in 0 of 5 runs, against 4 of 5 for a 36.5 GiB 2-bit alternative.
- Fewer slips in code. This model stays coherent in code generation.
- Rotor quantization of the routed experts (Gyro-S is built with APR, Agention Precision Rotor): the experts are stored in a rotated basis with a compact rotor code at 1.625โ1.875 bits per weight (1.71 on average), produced with our own encoder and calibration. The shared layers use a 5-bit K-quant encoded the same way.
- Tuned for coding, agentic work and long reasoning, with maths and 16 languages in the calibration.
Quick start (Docker)
Which setup?
| Your hardware | File and command | Notes |
|---|---|---|
| NVIDIA, 32 GB (RTX 5090) | Gyro-S, NVIDIA command below | The container runs it through Vulkan. Built from agentionai/llama.cpp with CUDA it decodes at ~100 tok/s with ~800 tok/s prefill (see Throughput). Up to 128k context (160k with -ub 256); the MTP draft fits up to 24k context (see Memory and context). |
| AMD, 32 GB (Radeon AI PRO R9700, W7800 32 GB) | Gyro-S, AMD command below | 58 tok/s decode, 1,246 tok/s prefill on an R9700. Same context limits; no MTP draft. |
| AMD Strix Halo, or other unified memory with 64 GB+ | Gyro-S or Gyro-M, AMD command | Let the GPU use at least 40 GiB for Gyro-S, 48 GiB for Gyro-M (BIOS UMA size / GTT limit). Add the MTP draft (below); up to 256k context with Gyro-S. |
| 48 GB of GPU memory or more (48 GB cards, 2ร24 GB) | Gyro-M | Higher quality (see Quality); add the MTP draft. |
| 16โ24 GB cards | Gyro-S with --n-cpu-moe N (the experts of N layers stay in system RAM, about 0.55 GiB each) |
Needs the container or agentionai/llama.cpp from 2026-10-02 or later. Start with N = 26 on 16 GB or N = 12 on 24 GB, lower it until the model just fits, and drop --mmproj to save ~1 GiB. Speed depends on your RAM bandwidth and CPU (AVX-512 helps most). Measured on Strix Halo: 15.5 tok/s with 23 layers on the CPU, 31 with none. Expect less on a desktop with dual-channel RAM. 32 GB of system RAM minimum, 64 GB recommended. |
Two GPUs? Put the whole model on one and the MTP draft on the other: --device Vulkan0 --device-draft Vulkan1.
Splitting the model's layers across both cards makes them take turns and is slower than one card (R9700: 57.8 tok/s on
one card, 37.6 split over two).
Tested on Linux with Vulkan (AMD Strix Halo, AMD R9700) and CUDA (NVIDIA RTX 5090 and RTX A6000, agentionai/llama.cpp main). ROCm, Windows and Apple (Metal) are untested or not supported yet.
Run it
The rotor code needs our llama.cpp build, agentionai/llama.cpp (Vulkan). The container has it ready:
hf download agentionai/Qwen3.8-Flash-Next-Gyro-GGUF Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf \
mtp-Qwen3.8-Flash-Next-draft.gguf mmproj-F16.gguf --local-dir ~/models
# AMD (and Intel) GPUs
docker run --rm -it --device /dev/dri --group-add "$(getent group render | cut -d: -f3)" \
-v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
-m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--mmproj /models/mmproj-F16.gguf
# NVIDIA GPUs: the same, with the NVIDIA container toolkit and its Vulkan driver
docker run --rm -it --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all \
-v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
-m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--mmproj /models/mmproj-F16.gguf
Then open http://localhost:8080. The sampling settings are Qwen's recommended thinking-mode parameters plus
--min-p 0.05, which trims the low-probability tokens that a low-bit quant lifts slightly.
--ngram-on-disk keeps the n-gram table off the GPU and out of RAM, and reads its rows with a fast parallel reader;
keep it in every command (plain memory-mapping, --lazy-mode on, roughly halves prompt-processing speed). Stock llama.cpp rejects this file as an
unknown type.
Building it yourself instead (needs Vulkan headers 1.4 or newer and glslc: on Ubuntu 22.04 the system packages
are too old, so install the LunarG Vulkan SDK and source its setup-env.sh
first):
git clone https://github.com/agentionai/llama.cpp && cd llama.cpp
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j # NVIDIA: -DGGML_CUDA=ON instead (CUDA toolkit 12.x+)
./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:Gyro-S -ngl 999 -fa on --jinja --ngram-on-disk \
-c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05
With -hf the vision projector (mmproj-F16.gguf) downloads and loads automatically, and so does the MTP draft
when you add --spec-type draft-mtp (below).
Vision: Gyro-S reads images through the vision projector mmproj-F16.gguf (0.9 GB, about 1 GiB of GPU
memory). It is on in the commands above: --mmproj /models/mmproj-F16.gguf with -m, automatic with -hf. To run
text-only and keep that memory, drop --mmproj (with -m) or add --no-mmproj (with -hf).
Faster with speculative decoding: add the MTP draft from this repo, mtp-Qwen3.8-Flash-Next-draft.gguf (2.7 GB,
about 3.8 GiB of GPU memory). On a single 32 GB card it fits next to Gyro-S at up to 24k context with
-ub 256 -ctkd q8_0 -ctvd q8_0 (32k with -ub 128, slower prompt processing); with 48 GB+ or unified memory there is
room at any context. Drafting helps most on code, JSON and editing; on fast cards (RTX 5090) it can slow down prose
and long reasoning, so try both. Flags:
--spec-type draft-mtp -md /models/mtp-Qwen3.8-Flash-Next-draft.gguf \
--spec-draft-n-max 6 --spec-draft-n-min 2 --spec-draft-mtp-vocab 32768
(With -hf, leave out -md: the draft is found by its mtp- name.)
Decode speed on Gyro-S (Strix Halo, greedy, tok/s): prose 32 โ 39, code 32 โ 52, JSON 32 โ 58, editing/copying
text 32 โ 61. --spec-draft-mtp-vocab needs the container image or agentionai/llama.cpp from 2026-10-01 on; older
builds ignore it with a warning. The draft costs about 3.8 GiB of extra GPU memory (see below).
Model overview
| Item | Specification |
|---|---|
| Base model | Qwen3.8-Flash-Next (architecture unchanged) |
| Routed experts | 48 layers ร 512 experts, 10 active per token. Rotor code: gate/up 1.625 bits/weight, down 1.875 |
| Shared layers | attention, linear attention and shared experts at 5-bit K-quant; router in BF16 |
| Weight basis | block-128 Hadamard rotation of each expert's input; the runtime applies the matching transform to activations |
| GPU memory | 27.65 GiB weights; 28.8 GiB total at 64k context |
| Download | 58.5 GB (includes the 26.8 GB n-gram table) |
| Backend | agentionai/llama.cpp, Vulkan (tested on AMD Strix Halo, gfx1151) |
Memory and context
What the GPU has to hold, as allocated by llama.cpp (weights + KV cache and recurrent state + compute buffers),
q8_0 KV cache, --ngram-on-disk, one slot:
| Context | GPU memory | + MTP draft (+3.75 GiB, measured) |
|---|---|---|
| 32k | 28.1 GiB | ~31.9 GiB |
| 64k | 28.8 GiB | ~32.6 GiB |
| 128k | 30.2 GiB | ~34.0 GiB |
| 256k | ~33.3 GiB (extrapolated) | ~37.1 GiB |
On a 32 GB card (about 31.5 GiB usable): up to 128k context without the MTP draft (160k with -ub 256). With the
draft: up to 24k context with -ub 256 -ctkd q8_0 -ctvd q8_0, or 32k with -ub 128 (measured on an RTX 5090); or put
the draft on a second GPU. The vision projector adds about 1 GiB; run text-only (--no-mmproj)
when memory is tight.
Short on memory? --n-cpu-moe N keeps the experts of the first N layers in system RAM (~0.55 GiB each); see
Which setup? for starting values. On Strix Halo, 23 layers on the CPU give 15.5 tok/s decode (31 with none), with
the CPU kernels in builds from 2026-10-02 on (2.35x faster than before).
Benchmarks
MMLU-Pro and GPQA Diamond results for Gyro-S and Gyro-M are being measured and will be added here.
Behaviour: a one-shot coding test
Same prompt (a voxel pagoda scene in one HTML file), same five seeds, reasoning effort medium, first answer only.
| File | GPU memory | Looped | Median reasoning |
|---|---|---|---|
| Gyro-S (this file) | 28.8 GiB | 0 / 5 | 23k tokens |
| ISTA-DASLab GSQ-RCO-IQ2_XS | 36.5 GiB | 4 / 5 | 28k tokens |
Quality
KL divergence and top-1 agreement against the unsloth Q8_0 source on a held-out corpus of 2026 technical writing
and code (-c 2048, 60 chunks). Wikitext is not used: this model's n-gram table has memorised it.
| File | GPU memory | KL div โ | top-1 agree โ |
|---|---|---|---|
| AP-Q4_K_XL | 67.4 GiB | 0.112 | 85.1 % |
| ISTA-DASLab GSQ-RCO-IQ2_XS | 36.5 GiB | 0.420 | 74.2 % |
| Gyro-M | 34.9 GiB | 0.306 | 78.1 % |
| Gyro-S | 27.65 GiB | 0.435 | 74.1 % |
KL divergence measures how closely the file follows the source token by token (error bars ยฑ0.003). Gyro-M's extra bits (2.125-bit experts, 8-bit shared layers) cut it by 30% against Gyro-S: median KLD 0.090 vs 0.141, and perplexity +3.1% above the source vs +8.2%. Gyro-M is ahead of the 36.5 GiB 2-bit alternative on both measures, at slightly less memory.
Throughput
llama-bench, batch size 1, no speculative decoding (tok/s):
| Hardware | File | Prefill (pp512) | Decode (tg128) |
|---|---|---|---|
| NVIDIA RTX 5090, 32 GB (CUDA, see below) | Gyro-S | 795 | 101 |
| NVIDIA RTX A6000, 48 GB (CUDA) | Gyro-S | 363 | 57.4 |
| NVIDIA RTX A6000, 48 GB (CUDA) | Gyro-M | 342 | 52.3 |
| AMD Radeon AI PRO R9700, 32 GB (one card) | Gyro-S | 1,246 | 57.8 |
| 2ร R9700, layers split across the cards | Gyro-S | 919 | 37.6 |
| AMD Strix Halo (Radeon 8060S), balanced power | Gyro-S | 250 | 31.7 |
| AMD Strix Halo (Radeon 8060S), balanced power | Gyro-M | 249 | 26.4 |
The R9700 numbers were measured by a tester with our Vulkan build, on a build with the same tensor types and sizes as Gyro-S. One card is faster than two: splitting layers makes the cards take turns. On Strix Halo, AP-Q4_K_XL reads 309 / 27.2 for comparison.
With MTP drafting on the R9700 (the earlier Q8_0 draft and flags, Qwen's sampling, thinking on): generating code 56.9 โ 70.6 tok/s (+24%), copying supplied code 57.5 โ 86.1 (+50%), explaining/debugging unchanged. Drafting pays most on predictable output; the current flags and draft (above) speed up the draft step itself.
NVIDIA RTX 5090 (32 GB), measured with our CUDA kernels (in agentionai/llama.cpp main; all 556 kernel tests pass on Blackwell), q8_0 KV cache, n-gram table on disk:
| Decode, by content (16k context, greedy) | prose 99, JSON 106, code 101, copied text 108 tok/s |
| Decode with context already filled | 118 (empty), 106 (8k), 91 (32k), 72 (64k) tok/s |
| GPU memory (of 31.8 GiB) | 28.7 GiB at 16k, 29.0 at 32k, 29.7 at 64k, 31.1 at 128k; 160k fits with -ub 256, 192k+ does not |
| MTP draft (cost-aware, 16k context) | JSON 172, code 138, copied text 187 tok/s; prose 98 (107 without drafting). Fits up to 24k context (-ub 256 -ctkd q8_0 -ctvd q8_0) or 32k (-ub 128) |
| Prefill | ~800 tok/s (pp512 795, pp2048 802); 905 on 2k-token prompts with -ub 2048 |
RTX 30-series / A-series (Ampere) with a CUDA build: the batched prompt path can crash (MUL_MAT_ID failed,
illegal memory access) when several requests with longer prompts run at once. Until the fix ships, start the server
with GGML_CUDA_TQ_MMQ=0 (or --parallel 1): decode speed is unchanged, prompt processing is slower.
NVIDIA with Vulkan (the current container): run Docker with --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all.
Many cloud GPU containers grant only compute,utility; then NVIDIA's Vulkan driver cannot start and llama.cpp
silently falls back to the CPU (no usable GPU found). Minimal images may also need libxext6 and libx11-6.
Status
Experimental. The format and kernels may change, and the file will be re-published when they do. Known limits: Vulkan and CUDA builds of agentionai/llama.cpp only; the n-gram table must stay on disk or in system RAM. Tested on AMD Strix Halo, Radeon AI PRO R9700, NVIDIA RTX 5090 and RTX A6000. We would like to hear how it runs on your card.
The weights come from our own encoder and calibration, which we are not publishing at this stage.
Support AgentionAI
These quants are released freely. If they save you VRAM or make Qwen more useful, you can buy me a coffee or some GPU time and sponsor continued quantization and benchmarking on GitHub. AgentionAI is a one-person team and can use your help.
- Downloads last month
- 3,081
Model tree for agentionai/Qwen3.8-Flash-Next-Gyro-GGUF
Base model
Qwen/Qwen3.8-Flash-Next