โš ๏ธ STOCK llama.cpp WILL NOT LOAD THIS MODEL

โš ๏ธ -fa off is required โ€” flash attention breaks the vision path on gfx1151.

14.11 GiB ยท 14.15 tok/s on a Ryzen AI MAX+ 395.

Phi-4-Reasoning-Vision-15B โ€” ROCmFPX 8-bit GGUF

An 8-bit ROCmFPX quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 GGUF โ€” a lossless source, not a requantization of a lower-bit build.

File Phi-4-reasoning-vision-15B-Q8_0_ROCMFPX.gguf
Size 14.11 GiB
BPW 8.27
ftype Q8_0_ROCMFPX (111)
mmproj mmproj-phi-4-reasoning-vision-15b-bf16.gguf (BF16, 862 MiB, included โ€” required for vision)

โ›” Requires a llama.cpp with the ROCmFPX quant types

Q8_0_ROCMFPX (111) and Q8_0_ROCMFPX_AGENT (115) exist only in charlie12345/ROCmFPX. Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.


All quant variants

Three builds, all measured in one session on one box with one binary (Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, ROCmFPX-2809dc5, -fa off) โ€” so these rows are directly comparable. Median of 3, warm-up discarded, otherwise-idle box.

variant ftype size bpw decode (median) range repo
4-bit COHERENT 102 7.93 GiB 4.65 24.91 24.88 โ€“ 24.91 link
8-bit AGENT 115 14.34 GiB 8.40 14.08 14.06 โ€“ 14.09 link
8-bit plain 111 14.11 GiB 8.27 14.15 14.13 โ€“ 14.20 link

โš ๏ธ The 4-bit build is ~1.7ร— faster and 43% smaller. The 8-bit builds exist for accuracy headroom, not throughput. The two 8-bit builds are within noise of each other (14.08 vs 14.15, ranges touching) โ€” this model has no MTP draft head, and AGENT's benefit shows up in draft acceptance, so there is nothing here for it to win.

โš ๏ธ -fa off is mandatory โ€” flash attention breaks the vision path on gfx1151. The BF16 mmproj (862 MiB) ships in every one of these repos and is required for vision.

Correctness: 17ร—23 โ‡’ โœ… 391 ยท capital of Japan โ‡’ โœ… Tokyo ยท days in 2024 โ‡’ โœ… 366

Give this model room to reason. On a curt "reply with only the number" prompt it can answer 365 for the 2024 question; allowed to reason it correctly derives leap year โ‡’ 366.


What was NOT measured

  • No perplexity run, and no quality A/B against the source. The checks above are memorized-fact prompts โ€” necessary but not sufficient.
  • No long-context testing. ยท No tool-calling evaluation.
  • Vision was smoke-tested only (the 4-bit build correctly describes an 8ร—8 red PNG as Red); no vision benchmark was run on these 8-bit builds.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
158
GGUF
Model size
15B params
Architecture
phi3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kingjones777/Phi-4-Reasoning-Vision-15B-ROCmFPX-Q8_0-GGUF

Base model

microsoft/phi-4
Quantized
(8)
this model