Qwen3.5-2B chat (text and images) on the AMD NPU
An IRON export of Qwen/Qwen3.5-2B for AMD Ryzen AI NPUs: the compiled NPU kernels (.xclbin + instruction streams) and the packed weights an IRON Rust runtime replays.
The kernels are compiled for NPU2 (AIE2P: Strix Point, Strix Halo, Krackan) and will not load on NPU1 (Phoenix, Hawk Point).
Download
hf download brishen/iron-qwen3.5-2b-npu2 --local-dir qwen3.5-2b
# or, from an IRON checkout:
python scripts/hf_models.py download qwen3.5-2b --repo brishen/iron-qwen3.5-2b-npu2 --out qwen3.5-2b
The text model and the vision tower (the multi-token prediction head is not exported). Every projection runs on the NPU: flm.GEMMs over the prompt and the image's patches, GEMVbfp16s for each generated token (15 tokens/s on a Ryzen AI 9 HX 370), and the 18 tokens/s; it needs taconite-qwen35 >= 0.1.1) is on the taconite-qwen35 Rust runtime replays it over XRT or directly over the amdxdna driver. ~2.5 GB. A 2.0 GB build with 6.5-bit text weights (bfp6s branch: hf download brishen/iron-qwen3.5-2b-npu2 --revision bfp6s --local-dir qwen3.5-2b. The previous bundle (~10 tokens/s): --revision bundle-w8.
Requirements
- An AMD Ryzen AI NPU2 (Strix Point, Strix Halo, Krackan) on Linux
with the
amdxdnadriver and its firmware (/dev/accel/accel0), and an unlimited locked-memory limit (ulimit -l unlimited): the weights live in ~2.5 GB of NPU-visible buffers. - XRT, unless the runtime is built with
--features direct(straight to the driver's ioctls, no XRT at all). - The model uses 3 of the NPU's 16 hardware-context slots (shared by all
processes). When other programs hold the rest, the runtime swaps its
contexts rather than failing (
--timingreports the swaps).
Usage
With the taconite-qwen35
runtime (docs):
cargo install taconite-qwen35 # over XRT
# or, with no XRT at all:
cargo install taconite-qwen35 --no-default-features --features cli,direct
qwen35 qwen3.5-2b --prompt "Why is the sky blue?" # streams the answer
qwen35 qwen3.5-2b --image photo.jpg --prompt "Describe this image." # --image repeats
qwen35 qwen3.5-2b --thinking --max-new 1024 --prompt "..." # reason in <think> first
qwen35 qwen3.5-2b --interactive # one prompt a line
qwen35 check qwen3.5-2b # verify this bundle
As a library:
use taconite_qwen35::{ChatOptions, Qwen35, RgbImage};
let mut q = Qwen35::load(std::path::Path::new("qwen3.5-2b"), 8192)?;
let opts = ChatOptions { images: vec![RgbImage { width, height, rgb }], ..Default::default() };
let (text, stats) = q.chat("Describe this image.", &opts, |s| print!("{s}"))?;
How it runs
Every weight matrix runs on the NPU; the host does the rest:
| NPU | host (f32) | |
|---|---|---|
| the prompt | every projection as an flm.GEMM over 256-row chunks, one hardware context |
tokenizer, embedding, norms, the Gated DeltaNet's causal conv and recurrence, partial M-RoPE, GQA attention |
| each generated token | every projection and the tied LM head as a GEMVbfp16, a second context |
the same, one row; greedy sampling |
| an image | the 24-block vision tower's projections as flm.GEMMs, a third context |
resize and patchify (torchvision-exact), LayerNorms, 2D RoPE, attention |
Weights are stored once, as bfp16 (8-bit mantissas, one exponent a block
of 8: 9 bits a weight): the decode GEMVs read the prefill GEMMs' packed
streams, and the input embedding is decoded from the tied LM head's. The
vision tower's weights are int8 per output channel, held exactly in bfp16
blocks with each channel's leftover scale factor applied on the host (a
single bfp16 block of the float weights loses too much over its 24
blocks). Activations enter the multiplies as bfp16 hi + lo (16 bits).
Accuracy
qwen35 check qwen3.5-2b on a Ryzen AI 9 HX 370, against transformers float32
on the ROCm iGPU, the next token teacher-forced on the reference's:
| prompt | prompt tokens | steps | top-1 agrees | disagreements (the reference's top-2 margin) | max |dlogit| (reference's top 16) |
|---|---|---|---|---|---|
| "Explain in three sentences why the sky is blue." | 23 | 97 | 94/97 | 3, all near-ties (<= 0.08) | 0.39 |
| a model-card summary | 778 | 110 | 108/110 | 2, all near-ties (<= 0.13) | 0.52 |
| a kitchen photo + "Describe this image in detail." | 280 | 160 | 154/160 | 6, all near-ties (<= 0.20) | 1.52 |
No disagreement where float32's top two are 0.5 or more apart. The
tokenizer and chat template reproduce Hugging Face's ids exactly, the
image preprocessing matches the Python reference bit for bit, and the
vision tower's output has cosine 0.992 against float32. The bundle
carries these references, so qwen35 check repeats the comparison on
your machine.
Speed
Ryzen AI 9 HX 370 (NPU2): setup 2 s; prompt 0.2 s for 23 tokens, 1.3 s
for 778; ~66 ms a generated token (15 tokens/s), bound by the ~2.1 GB of
weights a step streams from memory; an image of ~1M pixels (1040 patches)
~2-2.5 s through the vision tower (NPU ~0.25 s, host attention ~1.7 s).
Limits
- Greedy decoding only.
- A prompt of up to 2048 tokens (image tokens included); context up to
--max-ctx(default 8192). - Images are resized as the checkpoint's processor does, capped at ~1M pixels (4096 patches, 1024 image tokens); no video.
- The checkpoint's multi-token-prediction head is not used.
Files
manifest.txt (the kernels and model constants), tensors.txt /
tensors.bin (packed weights and references), kernels/ (xclbins and
instruction streams), tokenizer.json (the checkpoint's). Exported by
IRON's iron/applications/qwen3_5/export_qwen35.py.
Provenance
- Upstream model:
Qwen/Qwen3.5-2B(license:apache-2.0; its terms apply to these weights) - IRON commit:
bcf8c13 - Uploaded: 2026-09-30
- Files: 141, 2.5 GB