onwAMD / README_en.md
ryugyosoft's picture
onw AMD 0.1.1: FastFlowLM v1.0.7, NPU check reasons, diag.py
63211d2 verified
|
Raw History Blame Contribute Delete
3.14 kB
# onw AMD — Ore no NPU ga konna ni ugoku wake nai (AMD Ryzen AI NPU edition)
[日本語](README.md) | **English** ・ [Changelog](CHANGELOG_en.md)
An app that runs LLMs **on the AMD Ryzen AI NPU (XDNA2) alone**. It keeps the UI and the workings of
[onw](https://huggingface.co/ryugyosoft/onw) (the Intel NPU edition: tray app, model manager window, OpenAI-compatible API,
chat UI, check page, self-update) and replaces the inference with the NPU-only engine
[FastFlowLM](https://github.com/ROCm/FastFlowLM). How FastFlowLM is fetched, started and proxied is ported from
[Lemonade Server](https://github.com/lemonade-sdk/lemonade).
- **NPU only.** Nothing runs on the CPU, the GPU (iGPU) or as an NPU+GPU hybrid.
- The one engine is FastFlowLM (`flm` v1.0.7, about 40 MB), downloaded on first use. The models are the chat models
`flm list` offers (Qwen3.5 / 3.6, Gemma 4, gpt-oss, Llama 3.x, Phi-4-mini, LFM2, ...).
## Requirements
- **An AMD Ryzen AI PC with an XDNA2 NPU** (Ryzen AI 300 series [Strix Point / Krackan Point], Ryzen AI Max [Strix Halo] or newer).
XDNA1 (Ryzen 7040/8040, Phoenix / Hawk Point) is not supported by FastFlowLM.
- **The NPU driver** ("NPU Compute Accelerator Device" in the Device Manager; see
[Lemonade's driver guide](https://lemonade-server.ai/driver_install.html)); onw checks it with `flm validate`.
- Windows 11, or Linux (`amdxdna` driver)
- Disk: FastFlowLM about 40 MB + 1–22 GB per model
## Install
**Windows** (paste one line into PowerShell)
```powershell
irm https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.ps1 | iex
```
**Linux**
```bash
curl -fsSL https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.sh | bash
```
## Use
1. In the **Models** tab of the onw window, press "Install" next to FastFlowLM to list the models, then "Download" on
one (FastFlowLM is installed first when missing).
2. In the **Server** tab pick the model and press "Load". Once it says "Running", it is ready.
3. "Chat" opens the chat UI; the "Check page" measures time to first token and generation speed.
4. Other apps: **OpenAI-compatible API**, base URL `http://localhost:8010/v1`, model name as shown in the window
(e.g. `qwen3-it-4b-FLM`). Thinking (models labelled reasoning): `chat_template_kwargs: {"enable_thinking": true}` or
`reasoning_effort`; the thinking comes back as `reasoning_content`.
```bash
onw serve MODEL [--port 8010] [--host 0.0.0.0] [--api-key KEY] [--context 4096] [--open]
onw tray | onw window [--tab models] | onw list | onw npu
```
The server works like Lemonade Server: it starts `flm serve` as a child process on a free local port (8001 and up)
and forwards the requests with `model` set to FastFlowLM's checkpoint name. `/health` answers 200 once the model is loaded.
## License
Powered by [FastFlowLM](https://github.com/ROCm/FastFlowLM)
Code: Apache 2.0 (as onw). `onw/lemonade.py` ports code from Lemonade Server (Apache 2.0). FastFlowLM is not part of
onw; it is downloaded at run time from its own release (code MIT, NPU kernels closed-source but free for any use
including commercial). Models follow their own licenses.