Strix Halo GuideIndependent · measured · reproducible

AMD Ryzen AI MAX+ 395 · Radeon 8060S · 96GB/128GB

AMD Strix Halo local LLM setup, benchmarks and buyer guide.

Get from a retail Strix Halo machine to working local AI with tested BIOS and Linux settings, clear model choices, real Ollama and llama.cpp results, and buyer guidance that links every important claim to evidence.

  • Exact commands and versions
  • Positive and negative results
  • No pay-to-rank hardware

The short answer

Start simple. Move to advanced routes only when the workload needs them.

For most buyers, begin with Ubuntu 24.04, a small fixed UMA reservation, IOMMU enabled, Mesa/RADV and Ollama. Direct llama.cpp provides a separate benchmark route; ROCm, vLLM, MTP and performance forks stay workload-specific.

Beginner

Ollama + Vulkan/RADV

Ollama 0.31.2 is the reboot-qualified service baseline. Existing 0.32.15 passed scoped text/image/executed-tool and pinned-client checks after restart. Isolated 0.34.2 is useful but not default; official Qwen3.8 and upgrade/reboot acceptance remain open.

Read the scoped setup
Current model

Qwen3.8 27B

20.42 generation t/s and exact local retrieval through 50,059 prompt tokens on the measured official route.

Understand every route
Buyer

Choose memory before branding

Choose memory using the exact model artifact and required context. Compare complete configurations and dated regional offers; advertised capacity is not a local qualification.

Compare the systems

Evidence, not a collage

Useful results with their claim type still attached.

These numbers answer different questions. Direct llama-bench, an Ollama API result and a low-bit capacity pass are not interchangeable leaderboard rows.

Official Qwen3.8 buyer route

Chat, image, tools, thinking and medium context.

Nine warm Ollama 0.32.13 repeats measured 292.49 prompt t/s and 20.42 generation t/s. Exact retrieval passed through 50,059 prompt tokens.

API evidence on one 128GB Beelink; not direct llama-bench and not a 262K local claim.

See the context boundary
20.42t/sofficial Q4_K_M
warm generation mean

Balanced coding

96.76 t/s

Qwen3-Coder 30B-A3B

Direct llama-bench with the balanced UD-Q4_K_XL Vulkan/RADV route.

Inspect the run

Large-model capacity

284.33B

Ordinary GGUF on one box

A 90.86GB low-bit DeepSeek V4 Flash artifact loaded, generated at 13.27 tg128 and passed a narrow correctness smoke.

Capacity proof—not a broad model-quality recommendation.

Inspect the evidence

Community coverage

More than one chassis and one maintainer.

First-party Beelink measurements remain separate from community results on systems including Beelink, GMKtec, Corsair, Nimo and Minisforum. Framework is a buyer comparison, not a described owner-system reproduction in this evidence map. The count covers 10 described owner systems plus 3 external sources: 13 systems or independent sources, not 13 matched reproductions. Failures and limitations stay visible.

13
systems or independent sources
10
benchmark contributors

Why this guide exists

The differentiator is not another peak number. It is the route from purchase to repeatable result.

01

Buyers

Know what memory size, system and runtime fit the intended model.

02

Owners

Reach a working baseline without stitching together contradictory forum advice.

03

Reviewers

Use exact versions, commands, caveats and raw artifacts.

04

Vendors

See which firmware, driver and documentation gaps still block adoption.

Keep the evidence useful

Found a different result on your Strix Halo system?

Contributors keep public credit for their system, software stack, measurements and caveats. A correction or negative result is as useful as a faster row.

Independent by design.

Strix Halo Guide is an independent community project. It is not affiliated with, endorsed by, or an official publication of AMD or any OEM. Product names identify relevant or tested hardware—not partnerships.

Read the disclosure policy