# onw AMD — Ore no NPU ga konna ni ugoku wake nai (AMD Ryzen AI NPU edition) [日本語](README.md) | **English** ・ [Changelog](CHANGELOG_en.md) An app that runs LLMs **on the AMD Ryzen AI NPU (XDNA2) alone**. It keeps the UI and the workings of [onw](https://huggingface.co/ryugyosoft/onw) (the Intel NPU edition: tray app, model manager window, OpenAI-compatible API, chat UI, check page, self-update) and replaces the inference with the NPU-only engine [FastFlowLM](https://github.com/ROCm/FastFlowLM). How FastFlowLM is fetched, started and proxied is ported from [Lemonade Server](https://github.com/lemonade-sdk/lemonade). - **NPU only.** Nothing runs on the CPU, the GPU (iGPU) or as an NPU+GPU hybrid. - The one engine is FastFlowLM (`flm` v1.0.7, about 40 MB), downloaded on first use. The models are the chat models `flm list` offers (Qwen3.5 / 3.6, Gemma 4, gpt-oss, Llama 3.x, Phi-4-mini, LFM2, ...). ## Requirements - **An AMD Ryzen AI PC with an XDNA2 NPU** (Ryzen AI 300 series [Strix Point / Krackan Point], Ryzen AI Max [Strix Halo] or newer). XDNA1 (Ryzen 7040/8040, Phoenix / Hawk Point) is not supported by FastFlowLM. - **The NPU driver** ("NPU Compute Accelerator Device" in the Device Manager; see [Lemonade's driver guide](https://lemonade-server.ai/driver_install.html)); onw checks it with `flm validate`. - Windows 11, or Linux (`amdxdna` driver) - Disk: FastFlowLM about 40 MB + 1–22 GB per model ## Install **Windows** (paste one line into PowerShell) ```powershell irm https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.ps1 | iex ``` **Linux** ```bash curl -fsSL https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.sh | bash ``` ## Use 1. In the **Models** tab of the onw window, press "Install" next to FastFlowLM to list the models, then "Download" on one (FastFlowLM is installed first when missing). 2. In the **Server** tab pick the model and press "Load". Once it says "Running", it is ready. 3. "Chat" opens the chat UI; the "Check page" measures time to first token and generation speed. 4. Other apps: **OpenAI-compatible API**, base URL `http://localhost:8010/v1`, model name as shown in the window (e.g. `qwen3-it-4b-FLM`). Thinking (models labelled reasoning): `chat_template_kwargs: {"enable_thinking": true}` or `reasoning_effort`; the thinking comes back as `reasoning_content`. ```bash onw serve MODEL [--port 8010] [--host 0.0.0.0] [--api-key KEY] [--context 4096] [--open] onw tray | onw window [--tab models] | onw list | onw npu ``` The server works like Lemonade Server: it starts `flm serve` as a child process on a free local port (8001 and up) and forwards the requests with `model` set to FastFlowLM's checkpoint name. `/health` answers 200 once the model is loaded. ## License Powered by [FastFlowLM](https://github.com/ROCm/FastFlowLM) Code: Apache 2.0 (as onw). `onw/lemonade.py` ports code from Lemonade Server (Apache 2.0). FastFlowLM is not part of onw; it is downloaded at run time from its own release (code MIT, NPU kernels closed-source but free for any use including commercial). Models follow their own licenses.