onwAMD / CHANGELOG_en.md
ryugyosoft's picture
onw AMD 0.1.3: model file check, installer restart, download progress
00a5163 verified
|
Raw History Blame Contribute Delete
3.2 kB

Changelog

v0.1.3

  • Running the one-line installer again while onw is running (an update) now restarts the tray on the new version, loading the model again, instead of leaving the old tray running until a manual restart. Trays of 0.1.2 and older are stopped and the new tray takes over.
  • Model download progress comes from the bytes FastFlowLM reports and the model's real file sizes; it used to count finished files, so it sat at 0% while a big weights file came down (about 20 GB for qwen3.6-moe-35b-a3b).
  • Model files are checked against FastFlowLM's hashes: right after a download (damaged files are fetched again once), on demand from the Models tab (... -> "Check the files", then "Download again" for damaged ones), and by size before every load (broken weights give garbage such as <unused49>).
  • The model list no longer goes empty when flm list returns nothing because of one damaged model folder: it then comes from the model_list.json that ships with flm.
  • Fixed: installing FastFlowLM could set the chosen model to "engine".
  • Gemma 4's thinking (<|channel>thought) is split off too when the answer was cut off while thinking.
  • Windows: a process that did not get the single-instance lock no longer keeps holding it (it blocked the take-over).

v0.1.2

  • Linux: the installer (setup.sh) raises the memlock limit (required by FastFlowLM's Linux guide; Ubuntu's default fails flm validate) and asks for a restart when it did.
  • Models tab: how to fix a memlock failure, and a warning when another FastFlowLM is running (Lemonade Server, a system fastflowlm package's flm serve): they compete for the NPU.
  • diag.py also reports running FastFlowLM processes and installed fastflowlm / XRT / Lemonade packages.

v0.1.1

  • FastFlowLM updated to v1.0.7 (v1.0.6 failed every Linux kernel older than 6.17; it is replaced at the next download).
  • When FastFlowLM's NPU check (flm validate) fails, the Models tab says why (driver version, a link to the driver guide, "Check again", flm's own report); downloads stop with that reason.
  • New diag.py: NPU, driver, firmware, /dev/accel, memlock and FastFlowLM state in one file; --try-model runs a small model on the NPU to tell a wrong check from a real failure.

v0.1.0

First release of the AMD Ryzen AI NPU edition, based on onw (Intel NPU edition) v0.11.3.

  • Inference replaced with the NPU-only engine FastFlowLM v1.0.6 (fetching, starting and proxying ported from Lemonade Server), downloaded on first use.
  • Models: FastFlowLM's chat models. Nothing runs on the CPU, the GPU or as a hybrid.
  • Thinking switches mapped to FastFlowLM's think; the thinking comes back as reasoning_content.
  • Windows: FastFlowLM is tied to onw's server by a job object, so a crashed server never leaves it holding the NPU.
  • Detects the XDNA2 NPU and its driver; without one, downloads and loading stop and say why.
  • Intel-NPU-only settings removed (device, prompt lookup decoding, 64-token blocks, NPUW, expert memory).
  • The check page measures time to first token and generation speed by streaming.
  • The OpenVINO engine (graph building, KV cache, converter, measuring scripts) removed.