Qwen3.8-Flash-Next Gyro โ€” GGUF (experimental)

Qwen3.8-Flash-Next on a single GPU: one model, three sizes.

File For GPU memory (64k context) Status
Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf 32 GB cards 28.8 GiB available
Qwen3.8-Flash-Next-Gyro-M-TQ2_0.gguf 48 GB: 2ร—24 GB cards, 64 GB unified-memory machines ~37 GiB (34.35 GiB weights) available: KLD 0.306 (S: 0.435), top-1 78.1%
Qwen3.8-Flash-Next-Gyro-L.gguf 64 GB cards and unified-memory machines ~46โ€“50 GiB planned

All three keep the model's architecture; the larger files give the routed experts and the shared layers more bits. The sizes for unified-memory machines (such as Strix Halo) leave room for the operating system. All three need our llama.cpp build with Vulkan (see Quick start). The TQ1_0 / TQ2_0 in the file names is a size class for the Hub's file browser: the files use our own rotor formats, not the standard ternary types, and stock llama.cpp cannot load them.

The rest of this card describes Gyro-S: 28.8 GiB on the GPU at 64k context (our Q4: 67 GiB) | no loops in our looping test | with MTP drafting on one R9700 32 GB: 71 tok/s generating code, 86 tok/s editing supplied code (58 without drafting; 1,246 tok/s prefill). Its calibration set is weighted toward code and agentic work. This is an experimental release for people who want to try it and tell us how it behaves. The format and kernels may still change.

Highlights

  • Fits a 32 GB card. The mixture-of-experts model (126B parameters, 121B of them in the routed experts) needs 28.8 GiB of GPU memory with 64k of context (see Memory and context). Its 51.2B-parameter n-gram table stays on disk and is read row by row as needed.
  • Reasons without looping. Low-bit quants of this model tend to loop in long reasoning. In our test to induce looping (see Behaviour) this file looped in 0 of 5 runs, against 4 of 5 for a 36.5 GiB 2-bit alternative.
  • Fewer slips in code. This model stays coherent in code generation.
  • Rotor quantization of the routed experts (Gyro-S is built with APR, Agention Precision Rotor): the experts are stored in a rotated basis with a compact rotor code at 1.625โ€“1.875 bits per weight (1.71 on average), produced with our own encoder and calibration. The shared layers use a 5-bit K-quant encoded the same way.
  • Tuned for coding, agentic work and long reasoning, with maths and 16 languages in the calibration.

Quick start (Docker)

Which setup?

Your hardware File and command Notes
NVIDIA, 32 GB (RTX 5090) Gyro-S, NVIDIA command below The container runs it through Vulkan. Built from agentionai/llama.cpp with CUDA it decodes at ~100 tok/s with ~800 tok/s prefill (see Throughput). Up to 128k context (160k with -ub 256); the MTP draft fits up to 24k context (see Memory and context).
AMD, 32 GB (Radeon AI PRO R9700, W7800 32 GB) Gyro-S, AMD command below 58 tok/s decode, 1,246 tok/s prefill on an R9700. Same context limits; no MTP draft.
AMD Strix Halo, or other unified memory with 64 GB+ Gyro-S or Gyro-M, AMD command Let the GPU use at least 40 GiB for Gyro-S, 48 GiB for Gyro-M (BIOS UMA size / GTT limit). Add the MTP draft (below); up to 256k context with Gyro-S.
48 GB of GPU memory or more (48 GB cards, 2ร—24 GB) Gyro-M Higher quality (see Quality); add the MTP draft.
16โ€“24 GB cards Gyro-S with --n-cpu-moe N (the experts of N layers stay in system RAM, about 0.55 GiB each) Needs the container or agentionai/llama.cpp from 2026-10-02 or later. Start with N = 26 on 16 GB or N = 12 on 24 GB, lower it until the model just fits, and drop --mmproj to save ~1 GiB. Speed depends on your RAM bandwidth and CPU (AVX-512 helps most). Measured on Strix Halo: 15.5 tok/s with 23 layers on the CPU, 31 with none. Expect less on a desktop with dual-channel RAM. 32 GB of system RAM minimum, 64 GB recommended.

Two GPUs? Put the whole model on one and the MTP draft on the other: --device Vulkan0 --device-draft Vulkan1. Splitting the model's layers across both cards makes them take turns and is slower than one card (R9700: 57.8 tok/s on one card, 37.6 split over two).

Tested on Linux with Vulkan (AMD Strix Halo, AMD R9700) and CUDA (NVIDIA RTX 5090 and RTX A6000, agentionai/llama.cpp main). ROCm, Windows and Apple (Metal) are untested or not supported yet.

Run it

The rotor code needs our llama.cpp build, agentionai/llama.cpp (Vulkan). The container has it ready:

hf download agentionai/Qwen3.8-Flash-Next-Gyro-GGUF Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf \
  mtp-Qwen3.8-Flash-Next-draft.gguf mmproj-F16.gguf --local-dir ~/models

# AMD (and Intel) GPUs
docker run --rm -it --device /dev/dri --group-add "$(getent group render | cut -d: -f3)" \
  -v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
  -m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
  -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --mmproj /models/mmproj-F16.gguf

# NVIDIA GPUs: the same, with the NVIDIA container toolkit and its Vulkan driver
docker run --rm -it --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all \
  -v ~/models:/models -p 8080:8080 ghcr.io/agentionai/agention-llama:server \
  -m /models/Qwen3.8-Flash-Next-Gyro-S-TQ1_0.gguf -ngl 999 -fa on --jinja --ngram-on-disk \
  -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --mmproj /models/mmproj-F16.gguf

Then open http://localhost:8080. The sampling settings are Qwen's recommended thinking-mode parameters plus --min-p 0.05, which trims the low-probability tokens that a low-bit quant lifts slightly. --ngram-on-disk keeps the n-gram table off the GPU and out of RAM, and reads its rows with a fast parallel reader; keep it in every command (plain memory-mapping, --lazy-mode on, roughly halves prompt-processing speed). Stock llama.cpp rejects this file as an unknown type.

Building it yourself instead (needs Vulkan headers 1.4 or newer and glslc: on Ubuntu 22.04 the system packages are too old, so install the LunarG Vulkan SDK and source its setup-env.sh first):

git clone https://github.com/agentionai/llama.cpp && cd llama.cpp
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j      # NVIDIA: -DGGML_CUDA=ON instead (CUDA toolkit 12.x+)
./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-Gyro-GGUF:Gyro-S -ngl 999 -fa on --jinja --ngram-on-disk \
    -c 65536 -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05

With -hf the vision projector (mmproj-F16.gguf) downloads and loads automatically, and so does the MTP draft when you add --spec-type draft-mtp (below).

Vision: Gyro-S reads images through the vision projector mmproj-F16.gguf (0.9 GB, about 1 GiB of GPU memory). It is on in the commands above: --mmproj /models/mmproj-F16.gguf with -m, automatic with -hf. To run text-only and keep that memory, drop --mmproj (with -m) or add --no-mmproj (with -hf).

Faster with speculative decoding: add the MTP draft from this repo, mtp-Qwen3.8-Flash-Next-draft.gguf (2.7 GB, about 3.8 GiB of GPU memory). On a single 32 GB card it fits next to Gyro-S at up to 24k context with -ub 256 -ctkd q8_0 -ctvd q8_0 (32k with -ub 128, slower prompt processing); with 48 GB+ or unified memory there is room at any context. Drafting helps most on code, JSON and editing; on fast cards (RTX 5090) it can slow down prose and long reasoning, so try both. Flags:

--spec-type draft-mtp -md /models/mtp-Qwen3.8-Flash-Next-draft.gguf \
  --spec-draft-n-max 6 --spec-draft-n-min 2 --spec-draft-mtp-vocab 32768

(With -hf, leave out -md: the draft is found by its mtp- name.)

Decode speed on Gyro-S (Strix Halo, greedy, tok/s): prose 32 โ†’ 39, code 32 โ†’ 52, JSON 32 โ†’ 58, editing/copying text 32 โ†’ 61. --spec-draft-mtp-vocab needs the container image or agentionai/llama.cpp from 2026-10-01 on; older builds ignore it with a warning. The draft costs about 3.8 GiB of extra GPU memory (see below).

Model overview

Item Specification
Base model Qwen3.8-Flash-Next (architecture unchanged)
Routed experts 48 layers ร— 512 experts, 10 active per token. Rotor code: gate/up 1.625 bits/weight, down 1.875
Shared layers attention, linear attention and shared experts at 5-bit K-quant; router in BF16
Weight basis block-128 Hadamard rotation of each expert's input; the runtime applies the matching transform to activations
GPU memory 27.65 GiB weights; 28.8 GiB total at 64k context
Download 58.5 GB (includes the 26.8 GB n-gram table)
Backend agentionai/llama.cpp, Vulkan (tested on AMD Strix Halo, gfx1151)

Memory and context

What the GPU has to hold, as allocated by llama.cpp (weights + KV cache and recurrent state + compute buffers), q8_0 KV cache, --ngram-on-disk, one slot:

Context GPU memory + MTP draft (+3.75 GiB, measured)
32k 28.1 GiB ~31.9 GiB
64k 28.8 GiB ~32.6 GiB
128k 30.2 GiB ~34.0 GiB
256k ~33.3 GiB (extrapolated) ~37.1 GiB

On a 32 GB card (about 31.5 GiB usable): up to 128k context without the MTP draft (160k with -ub 256). With the draft: up to 24k context with -ub 256 -ctkd q8_0 -ctvd q8_0, or 32k with -ub 128 (measured on an RTX 5090); or put the draft on a second GPU. The vision projector adds about 1 GiB; run text-only (--no-mmproj) when memory is tight. Short on memory? --n-cpu-moe N keeps the experts of the first N layers in system RAM (~0.55 GiB each); see Which setup? for starting values. On Strix Halo, 23 layers on the CPU give 15.5 tok/s decode (31 with none), with the CPU kernels in builds from 2026-10-02 on (2.35x faster than before).

Benchmarks

MMLU-Pro and GPQA Diamond results for Gyro-S and Gyro-M are being measured and will be added here.

Behaviour: a one-shot coding test

Same prompt (a voxel pagoda scene in one HTML file), same five seeds, reasoning effort medium, first answer only.

File GPU memory Looped Median reasoning
Gyro-S (this file) 28.8 GiB 0 / 5 23k tokens
ISTA-DASLab GSQ-RCO-IQ2_XS 36.5 GiB 4 / 5 28k tokens

Quality

KL divergence and top-1 agreement against the unsloth Q8_0 source on a held-out corpus of 2026 technical writing and code (-c 2048, 60 chunks). Wikitext is not used: this model's n-gram table has memorised it.

File GPU memory KL div โ†“ top-1 agree โ†‘
AP-Q4_K_XL 67.4 GiB 0.112 85.1 %
ISTA-DASLab GSQ-RCO-IQ2_XS 36.5 GiB 0.420 74.2 %
Gyro-M 34.9 GiB 0.306 78.1 %
Gyro-S 27.65 GiB 0.435 74.1 %

KL divergence measures how closely the file follows the source token by token (error bars ยฑ0.003). Gyro-M's extra bits (2.125-bit experts, 8-bit shared layers) cut it by 30% against Gyro-S: median KLD 0.090 vs 0.141, and perplexity +3.1% above the source vs +8.2%. Gyro-M is ahead of the 36.5 GiB 2-bit alternative on both measures, at slightly less memory.

Throughput

llama-bench, batch size 1, no speculative decoding (tok/s):

Hardware File Prefill (pp512) Decode (tg128)
NVIDIA RTX 5090, 32 GB (CUDA, see below) Gyro-S 795 101
NVIDIA RTX A6000, 48 GB (CUDA) Gyro-S 363 57.4
NVIDIA RTX A6000, 48 GB (CUDA) Gyro-M 342 52.3
AMD Radeon AI PRO R9700, 32 GB (one card) Gyro-S 1,246 57.8
2ร— R9700, layers split across the cards Gyro-S 919 37.6
AMD Strix Halo (Radeon 8060S), balanced power Gyro-S 250 31.7
AMD Strix Halo (Radeon 8060S), balanced power Gyro-M 249 26.4

The R9700 numbers were measured by a tester with our Vulkan build, on a build with the same tensor types and sizes as Gyro-S. One card is faster than two: splitting layers makes the cards take turns. On Strix Halo, AP-Q4_K_XL reads 309 / 27.2 for comparison.

With MTP drafting on the R9700 (the earlier Q8_0 draft and flags, Qwen's sampling, thinking on): generating code 56.9 โ†’ 70.6 tok/s (+24%), copying supplied code 57.5 โ†’ 86.1 (+50%), explaining/debugging unchanged. Drafting pays most on predictable output; the current flags and draft (above) speed up the draft step itself.

NVIDIA RTX 5090 (32 GB), measured with our CUDA kernels (in agentionai/llama.cpp main; all 556 kernel tests pass on Blackwell), q8_0 KV cache, n-gram table on disk:

Decode, by content (16k context, greedy) prose 99, JSON 106, code 101, copied text 108 tok/s
Decode with context already filled 118 (empty), 106 (8k), 91 (32k), 72 (64k) tok/s
GPU memory (of 31.8 GiB) 28.7 GiB at 16k, 29.0 at 32k, 29.7 at 64k, 31.1 at 128k; 160k fits with -ub 256, 192k+ does not
MTP draft (cost-aware, 16k context) JSON 172, code 138, copied text 187 tok/s; prose 98 (107 without drafting). Fits up to 24k context (-ub 256 -ctkd q8_0 -ctvd q8_0) or 32k (-ub 128)
Prefill ~800 tok/s (pp512 795, pp2048 802); 905 on 2k-token prompts with -ub 2048

RTX 30-series / A-series (Ampere) with a CUDA build: the batched prompt path can crash (MUL_MAT_ID failed, illegal memory access) when several requests with longer prompts run at once. Until the fix ships, start the server with GGML_CUDA_TQ_MMQ=0 (or --parallel 1): decode speed is unchanged, prompt processing is slower.

NVIDIA with Vulkan (the current container): run Docker with --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all. Many cloud GPU containers grant only compute,utility; then NVIDIA's Vulkan driver cannot start and llama.cpp silently falls back to the CPU (no usable GPU found). Minimal images may also need libxext6 and libx11-6.

Status

Experimental. The format and kernels may change, and the file will be re-published when they do. Known limits: Vulkan and CUDA builds of agentionai/llama.cpp only; the n-gram table must stay on disk or in system RAM. Tested on AMD Strix Halo, Radeon AI PRO R9700, NVIDIA RTX 5090 and RTX A6000. We would like to hear how it runs on your card.

The weights come from our own encoder and calibration, which we are not publishing at this stage.

Support AgentionAI

These quants are released freely. If they save you VRAM or make Qwen more useful, you can buy me a coffee or some GPU time and sponsor continued quantization and benchmarking on GitHub. AgentionAI is a one-person team and can use your help.

Downloads last month
3,081
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for agentionai/Qwen3.8-Flash-Next-Gyro-GGUF

Quantized
(325)
this model

Space using agentionai/Qwen3.8-Flash-Next-Gyro-GGUF 1