Qwen3.5-2B chat (text and images) on the AMD NPU

An IRON export of Qwen/Qwen3.5-2B for AMD Ryzen AI NPUs: the compiled NPU kernels (.xclbin + instruction streams) and the packed weights an IRON Rust runtime replays.

The kernels are compiled for NPU2 (AIE2P: Strix Point, Strix Halo, Krackan) and will not load on NPU1 (Phoenix, Hawk Point).

Download

hf download brishen/iron-qwen3.5-2b-npu2 --local-dir qwen3.5-2b
# or, from an IRON checkout:
python scripts/hf_models.py download qwen3.5-2b --repo brishen/iron-qwen3.5-2b-npu2 --out qwen3.5-2b

The text model and the vision tower (the multi-token prediction head is not exported). Every projection runs on the NPU: flm.GEMMs over the prompt and the image's patches, GEMVbfp16s for each generated token (15 tokens/s on a Ryzen AI 9 HX 370), and the taconite-qwen35 Rust runtime replays it over XRT or directly over the amdxdna driver. ~2.5 GB. A 2.0 GB build with 6.5-bit text weights (18 tokens/s; it needs taconite-qwen35 >= 0.1.1) is on the bfp6s branch: hf download brishen/iron-qwen3.5-2b-npu2 --revision bfp6s --local-dir qwen3.5-2b. The previous bundle (~10 tokens/s): --revision bundle-w8.

Requirements

  • An AMD Ryzen AI NPU2 (Strix Point, Strix Halo, Krackan) on Linux with the amdxdna driver and its firmware (/dev/accel/accel0), and an unlimited locked-memory limit (ulimit -l unlimited): the weights live in ~2.5 GB of NPU-visible buffers.
  • XRT, unless the runtime is built with --features direct (straight to the driver's ioctls, no XRT at all).
  • The model uses 3 of the NPU's 16 hardware-context slots (shared by all processes). When other programs hold the rest, the runtime swaps its contexts rather than failing (--timing reports the swaps).

Usage

With the taconite-qwen35 runtime (docs):

cargo install taconite-qwen35   # over XRT
# or, with no XRT at all:
cargo install taconite-qwen35 --no-default-features --features cli,direct

qwen35 qwen3.5-2b --prompt "Why is the sky blue?"                       # streams the answer
qwen35 qwen3.5-2b --image photo.jpg --prompt "Describe this image."    # --image repeats
qwen35 qwen3.5-2b --thinking --max-new 1024 --prompt "..."             # reason in <think> first
qwen35 qwen3.5-2b --interactive                                        # one prompt a line
qwen35 check qwen3.5-2b                                                # verify this bundle

As a library:

use taconite_qwen35::{ChatOptions, Qwen35, RgbImage};

let mut q = Qwen35::load(std::path::Path::new("qwen3.5-2b"), 8192)?;
let opts = ChatOptions { images: vec![RgbImage { width, height, rgb }], ..Default::default() };
let (text, stats) = q.chat("Describe this image.", &opts, |s| print!("{s}"))?;

How it runs

Every weight matrix runs on the NPU; the host does the rest:

NPU host (f32)
the prompt every projection as an flm.GEMM over 256-row chunks, one hardware context tokenizer, embedding, norms, the Gated DeltaNet's causal conv and recurrence, partial M-RoPE, GQA attention
each generated token every projection and the tied LM head as a GEMVbfp16, a second context the same, one row; greedy sampling
an image the 24-block vision tower's projections as flm.GEMMs, a third context resize and patchify (torchvision-exact), LayerNorms, 2D RoPE, attention

Weights are stored once, as bfp16 (8-bit mantissas, one exponent a block of 8: 9 bits a weight): the decode GEMVs read the prefill GEMMs' packed streams, and the input embedding is decoded from the tied LM head's. The vision tower's weights are int8 per output channel, held exactly in bfp16 blocks with each channel's leftover scale factor applied on the host (a single bfp16 block of the float weights loses too much over its 24 blocks). Activations enter the multiplies as bfp16 hi + lo (16 bits).

Accuracy

qwen35 check qwen3.5-2b on a Ryzen AI 9 HX 370, against transformers float32 on the ROCm iGPU, the next token teacher-forced on the reference's:

prompt prompt tokens steps top-1 agrees disagreements (the reference's top-2 margin) max |dlogit| (reference's top 16)
"Explain in three sentences why the sky is blue." 23 97 94/97 3, all near-ties (<= 0.08) 0.39
a model-card summary 778 110 108/110 2, all near-ties (<= 0.13) 0.52
a kitchen photo + "Describe this image in detail." 280 160 154/160 6, all near-ties (<= 0.20) 1.52

No disagreement where float32's top two are 0.5 or more apart. The tokenizer and chat template reproduce Hugging Face's ids exactly, the image preprocessing matches the Python reference bit for bit, and the vision tower's output has cosine 0.992 against float32. The bundle carries these references, so qwen35 check repeats the comparison on your machine.

Speed

Ryzen AI 9 HX 370 (NPU2): setup 2 s; prompt 0.2 s for 23 tokens, 1.3 s for 778; ~66 ms a generated token (15 tokens/s), bound by the ~2.1 GB of weights a step streams from memory; an image of ~1M pixels (1040 patches) ~2-2.5 s through the vision tower (NPU ~0.25 s, host attention ~1.7 s).

Limits

  • Greedy decoding only.
  • A prompt of up to 2048 tokens (image tokens included); context up to --max-ctx (default 8192).
  • Images are resized as the checkpoint's processor does, capped at ~1M pixels (4096 patches, 1024 image tokens); no video.
  • The checkpoint's multi-token-prediction head is not used.

Files

manifest.txt (the kernels and model constants), tensors.txt / tensors.bin (packed weights and references), kernels/ (xclbins and instruction streams), tokenizer.json (the checkpoint's). Exported by IRON's iron/applications/qwen3_5/export_qwen35.py.

Provenance

  • Upstream model: Qwen/Qwen3.5-2B (license: apache-2.0; its terms apply to these weights)
  • IRON commit: bcf8c13
  • Uploaded: 2026-09-30
  • Files: 141, 2.5 GB
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for brishen/iron-qwen3.5-2b-npu2

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(447)
this model