Independent guide · methods and evidence

Qwen3.8 27B on AMD Strix Halo

Run Qwen3.8 27B on AMD Strix Halo: measured Ollama performance, image and tool checks, context limits, and separately labelled experimental routes.

Evidence reviewed .

September 19 functional update: the scoped runtime qualification records text (Qwen3.6), image (Qwen2.5-VL), executed tools (Devstral) and pinned Open WebUI on the existing 0.32.15 service after restart. Isolated 0.34.2 is useful but not default; official Qwen3.8 and full-reboot candidate acceptance remain open. The tested host's Ollama listener was LAN-reachable, not local-only. Released llama.cpp v0.4.1 passed bounded direct/server/HIP controls; this does not qualify every model, long-context shape or maximum-memory allocation.

Qwen3.8 27B runs locally on AMD Ryzen AI MAX+ 395 / Radeon 8060S Strix Halo systems. The useful question is no longer only “does it run?” It is which official, stock, MTP, DFlash, ROCmFP4, or performance-fork route fits the workload—and which published numbers are actually comparable.

Evidence reviewed: September 19, 2026.

Fast Answer

Question Current answer
Easiest measured official route qwen3.8:27b through Ollama 0.32.13 and Vulkan/RADV
Guide-measured warm result 292.49 prompt t/s and 20.42 generation t/s over nine repeats
Guide-measured capabilities Image, tool-call, thinking, and exact retrieval through 50,059 prompt tokens passed
Long-context boundary A 56,051-token attempt caused a recoverable device loss on that exact stack; separate corrected GMKtec evidence reached 261,130 evaluated tokens
52-65 t/s posts Advanced fork/quant/speculation leads that require their exact artifacts, prompts, context behavior, and independent reproduction

For commands, the complete matrix, caveats, upstream alerts, and raw evidence, read the canonical Qwen3.8 Strix Halo decision page.

Start With Ollama

The measured official route used the 17.7GB-decimal Q4_K_M artifact:

ollama run qwen3.8:27b

The Strix Halo service still needs the guide's Vulkan/iGPU environment, including OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1. Ollama 0.34.2 is the current checked package, but it has not inherited the measured 0.32.13 result or the full normal-service/reboot qualification.

Why The Speed Claims Differ

Official Ollama, stock llama.cpp, custom quants, ROCmFP4, native MTP, DFlash2, and adaptive speculation are different routes. Code versus prose, cold versus cached context, context depth, generated-token count, and draft acceptance can change the result again. A useful comparison therefore records the entire profile, not only a tokens-per-second screenshot.

The guide keeps:

separate so readers can decide which result applies to them.

What Should Be Tested Next?

The missing proof is a matched same-host ladder: stock b10687 Vulkan without speculation, native MTP, corrected HIP correctness, a published ROCmFP4 route, and a fully published DFlash/adaptive route. It must cover code and prose, 4K/16K/50K context, exact outputs, vision/tools, acceptance, memory, hashes, and raw logs.

See the live test queue or submit a benchmark report.

Independence And Affiliate Disclosure

This guide contains no affiliate links as of September 19, 2026. Future affiliate, loaned, gifted, sponsored, or early-access relationships must be disclosed near the relevant links/results and do not buy positive conclusions. Community results remain separate from first-party measurements. The public affiliate link registry remains the audit source if that status changes.

Independent by design.

Strix Halo Guide is an independent community project. It is not affiliated with, endorsed by, or an official publication of AMD or any OEM. Product names identify relevant or tested hardware—not partnerships.

Read the disclosure policy