qwen3-embedding-0.6b for the Ryzen AI NPU
Qwen3-Embedding-0.6B, converted to run on the Ryzen AI NPU of AMD Strix Halo inside halogen-flash-server, beside the Flash model on the GPU and behind the same port. Questions go to the Discord. Bugs go to the server repository's issues.
These files run only inside the halogen-flash-server container, on the NPU. They will not load in transformers or any other runtime.
What is here
devices/ 62 MB the NPU program for this model, and its NOTICE.md
qwen3-embedding-0.6b.hnpw 813 MB the weights, converted for the NPU
tokenizer/ 11 MB the tokenizer
Since 0.16.2 the reranker (halogen-npu-qwen3-reranker-0.6b) and the moderation model (halogen-npu-qwen3guard-gen-0.6b) run on this NPU program too, so a host that serves more than one fetches it once.
Use it
Start halogen-flash-server 0.16.0 or later with HALOGEN_DOWNLOAD on and
HALOGEN_NPU_MODELS=qwen3-embedding-0.6b. The server fetches these files into
/models/npu/qwen3-embedding-0.6b/ and checks each one's size and checksum before it starts.
It answers /v1/embeddings with 1,024-dimension vectors (fewer with dimensions), in the OpenAI request shape.
The host needs the NPU driver, XRT with its NPU plugin, and the GPU's fabric clock held while the NPU works beside the GPU. The setup, every request shape, and how to run your own fine-tune of this model are in docs/NPU.md.
Source and license
Converted from Qwen/Qwen3-Embedding-0.6B at revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3. Apache-2.0, as the source.
devices/NOTICE.md lists the third-party notices the NPU program carries.