halogen

qwen3-embedding-0.6b for the Ryzen AI NPU

Qwen3-Embedding-0.6B, converted to run on the Ryzen AI NPU of AMD Strix Halo inside halogen-flash-server, beside the Flash model on the GPU and behind the same port. Questions go to the Discord. Bugs go to the server repository's issues.

These files run only inside the halogen-flash-server container, on the NPU. They will not load in transformers or any other runtime.

What is here

devices/              62 MB  the NPU program for this model, and its NOTICE.md
qwen3-embedding-0.6b.hnpw    813 MB  the weights, converted for the NPU
tokenizer/            11 MB  the tokenizer

Since 0.16.2 the reranker (halogen-npu-qwen3-reranker-0.6b) and the moderation model (halogen-npu-qwen3guard-gen-0.6b) run on this NPU program too, so a host that serves more than one fetches it once.

Use it

Start halogen-flash-server 0.16.0 or later with HALOGEN_DOWNLOAD on and HALOGEN_NPU_MODELS=qwen3-embedding-0.6b. The server fetches these files into /models/npu/qwen3-embedding-0.6b/ and checks each one's size and checksum before it starts. It answers /v1/embeddings with 1,024-dimension vectors (fewer with dimensions), in the OpenAI request shape.

The host needs the NPU driver, XRT with its NPU plugin, and the GPU's fabric clock held while the NPU works beside the GPU. The setup, every request shape, and how to run your own fine-tune of this model are in docs/NPU.md.

Source and license

Converted from Qwen/Qwen3-Embedding-0.6B at revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3. Apache-2.0, as the source. devices/NOTICE.md lists the third-party notices the NPU program carries.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for peonist-ai/halogen-npu-qwen3-embedding-0.6b

Quantized
(263)
this model