onwAMD / README_en.md
ryugyosoft's picture
onw AMD 0.1.1: FastFlowLM v1.0.7, NPU check reasons, diag.py
63211d2 verified
|
Raw History Blame Contribute Delete
3.14 kB

onw AMD — Ore no NPU ga konna ni ugoku wake nai (AMD Ryzen AI NPU edition)

日本語 | English ・ Changelog

An app that runs LLMs on the AMD Ryzen AI NPU (XDNA2) alone. It keeps the UI and the workings of onw (the Intel NPU edition: tray app, model manager window, OpenAI-compatible API, chat UI, check page, self-update) and replaces the inference with the NPU-only engine FastFlowLM. How FastFlowLM is fetched, started and proxied is ported from Lemonade Server.

  • NPU only. Nothing runs on the CPU, the GPU (iGPU) or as an NPU+GPU hybrid.
  • The one engine is FastFlowLM (flm v1.0.7, about 40 MB), downloaded on first use. The models are the chat models flm list offers (Qwen3.5 / 3.6, Gemma 4, gpt-oss, Llama 3.x, Phi-4-mini, LFM2, ...).

Requirements

  • An AMD Ryzen AI PC with an XDNA2 NPU (Ryzen AI 300 series [Strix Point / Krackan Point], Ryzen AI Max [Strix Halo] or newer). XDNA1 (Ryzen 7040/8040, Phoenix / Hawk Point) is not supported by FastFlowLM.
  • The NPU driver ("NPU Compute Accelerator Device" in the Device Manager; see Lemonade's driver guide); onw checks it with flm validate.
  • Windows 11, or Linux (amdxdna driver)
  • Disk: FastFlowLM about 40 MB + 1–22 GB per model

Install

Windows (paste one line into PowerShell)

irm https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.ps1 | iex

Linux

curl -fsSL https://huggingface.co/ryugyosoft/onwAMD/resolve/main/install.sh | bash

Use

  1. In the Models tab of the onw window, press "Install" next to FastFlowLM to list the models, then "Download" on one (FastFlowLM is installed first when missing).
  2. In the Server tab pick the model and press "Load". Once it says "Running", it is ready.
  3. "Chat" opens the chat UI; the "Check page" measures time to first token and generation speed.
  4. Other apps: OpenAI-compatible API, base URL http://localhost:8010/v1, model name as shown in the window (e.g. qwen3-it-4b-FLM). Thinking (models labelled reasoning): chat_template_kwargs: {"enable_thinking": true} or reasoning_effort; the thinking comes back as reasoning_content.
onw serve MODEL [--port 8010] [--host 0.0.0.0] [--api-key KEY] [--context 4096] [--open]
onw tray | onw window [--tab models] | onw list | onw npu

The server works like Lemonade Server: it starts flm serve as a child process on a free local port (8001 and up) and forwards the requests with model set to FastFlowLM's checkpoint name. /health answers 200 once the model is loaded.

License

Powered by FastFlowLM

Code: Apache 2.0 (as onw). onw/lemonade.py ports code from Lemonade Server (Apache 2.0). FastFlowLM is not part of onw; it is downloaded at run time from its own release (code MIT, NPU kernels closed-source but free for any use including commercial). Models follow their own licenses.