Direct · API · server · community · capacity

AMD Strix Halo LLM benchmarks you can inspect.

Every promoted result keeps its system, backend, runtime, build, model, quantization, workload, structured row, raw evidence and caveat attached. Different claim types stay separate.

Evidence reviewed .

Read the numbers correctly

Four labels prevent false comparisons.

Direct benchmark

Controlled llama-bench rows with an exact model, quant, build and command shape.

API or server

Ollama, llama-server, batching and speculative results include service behavior that direct rows do not.

Community evidence

Credited external systems stay separate until independently reproduced and normalized.

Capacity proof

A large file loading and generating proves fit and basic operation—not broad quality or speed.

Selected headline evidence

A route map, not one misleading leaderboard.

QuestionRouteResultEvidence class
Current official multimodal chatQwen3.8 27B Q4_K_M · Ollama 0.32.13 Vulkan292.49 prompt / 20.42 generation t/sFirst-party API
Balanced codingQwen3-Coder 30B-A3B UD-Q4_K_XL · direct Vulkan96.76 tg128First-party direct
Speed-first 30B codingQwen3-Coder 30B-A3B Q4_K_S · direct Vulkan100.99 tg128First-party direct
Small-MoE speedLFM2.5 8B-A1B Q4_K_M · direct Vulkan168.96 tg128First-party direct
80B MoE routeQwen3-Next 80B-A3B UD-Q4_K_XL · direct Vulkan59.06 tg128First-party direct
120B-class one-box capacityNemotron 3 Super 120B-A12B UD-IQ4_XS18.43 tg128Capacity/direct
284B ordinary-GGUF capacityDeepSeek V4 Flash UD-IQ2_XXS13.27 tg128Capacity/basic correctness

Prompt processing and generation are different metrics. Server aggregate throughput is not single-stream decode. Values from different rows are not normalized into a universal OEM or model ranking.

Evidence coverage

The guide now represents 13 systems or independent sources and 10 benchmark contributors.

The auditable split is 10 described owner systems plus 3 independently attributable external sources. Repeated evidence from one physical machine counts once; separate machines can count separately even when they share a product model or owner. External packages remain labelled as sources when unique hardware identity cannot be safely proved.

Coverage includes Linux and Windows, thermals, power, multi-user serving, RPC, NPU sidecars, long context, vision and failed routes.

13
systems or sources
10
contributors
300
GitHub stars in the 2026-08-27 snapshot

Claim-to-proof chain

How to verify any number.

  1. Read the claim row for system, backend, build, model, quant and workload.
  2. Open the structured CSV for the exact measurement and repeats.
  3. Inspect the raw directory for commands, logs, host state and failures.
  4. Read the caveat before transferring the result to a different system or workload.

Strengthen the map

A slower or failed result can remove more buyer uncertainty than another peak.

Submit the exact system, firmware, runtime, command, model, quant, repeats and raw output. Community attribution and caveats stay attached.

August 30: direct sentinel and Flash-Next scout

These first-party direct llama-bench rows use locally built b10687 (c841aee), kernel 7.0.0-30, Mesa/RADV 26.1.7 and the desktop performance profile. AMDGPU DPM stayed on auto; the recorded CPU-only background workload remained active. They are current-stack compatibility/speed observations, not replacements for strict-clean headline runs or controlled comparisons against earlier kernels.

Model Quant pp512 t/s tg128 t/s Repeats Class
Qwen3-Coder 30B-A3B UD-Q4_K_XL 1264.16 94.64 20 Sentinel
Qwen3-Next 80B-A3B UD-Q4_K_XL 675.76 62.09 20 Sentinel
Qwen3.8-Flash-Next UD-IQ4_XS 394.73 27.16 10 Single-artifact scout

The Flash-Next GGUF is approximately 93.7GB. Its separate arithmetic smoke returned 56 for 7×8 after the power profile was restored. Cite the repeated benchmark for speed, not that smoke. Long context, vision, tools, server behavior and broad task quality are not qualified. The official model card counts 125B/6B active plus 51B n-gram embeddings and 4B MTP separately; the tested GGUF reports about 177B parameters. It uses qwen-community-1.0, not Apache 2.0; check the model terms before commercial deployment.

Evidence: structured rows, claim registry, sentinel notes and raw CSVs, Flash-Next notes and raw CSV, artifact hashes.

Independent by design.

Strix Halo Guide is an independent community project. It is not affiliated with, endorsed by, or an official publication of AMD or any OEM. Product names identify relevant or tested hardware—not partnerships.

Read the disclosure policy