Direct benchmark
Controlled llama-bench rows with an exact model, quant, build and command shape.
Direct · API · server · community · capacity
Every promoted result keeps its system, backend, runtime, build, model, quantization, workload, structured row, raw evidence and caveat attached. Different claim types stay separate.
Evidence reviewed .
Read the numbers correctly
Controlled llama-bench rows with an exact model, quant, build and command shape.
Ollama, llama-server, batching and speculative results include service behavior that direct rows do not.
Credited external systems stay separate until independently reproduced and normalized.
A large file loading and generating proves fit and basic operation—not broad quality or speed.
Selected headline evidence
| Question | Route | Result | Evidence class |
|---|---|---|---|
| Current official multimodal chat | Qwen3.8 27B Q4_K_M · Ollama 0.32.13 Vulkan | 292.49 prompt / 20.42 generation t/s | First-party API |
| Balanced coding | Qwen3-Coder 30B-A3B UD-Q4_K_XL · direct Vulkan | 96.76 tg128 | First-party direct |
| Speed-first 30B coding | Qwen3-Coder 30B-A3B Q4_K_S · direct Vulkan | 100.99 tg128 | First-party direct |
| Small-MoE speed | LFM2.5 8B-A1B Q4_K_M · direct Vulkan | 168.96 tg128 | First-party direct |
| 80B MoE route | Qwen3-Next 80B-A3B UD-Q4_K_XL · direct Vulkan | 59.06 tg128 | First-party direct |
| 120B-class one-box capacity | Nemotron 3 Super 120B-A12B UD-IQ4_XS | 18.43 tg128 | Capacity/direct |
| 284B ordinary-GGUF capacity | DeepSeek V4 Flash UD-IQ2_XXS | 13.27 tg128 | Capacity/basic correctness |
Prompt processing and generation are different metrics. Server aggregate throughput is not single-stream decode. Values from different rows are not normalized into a universal OEM or model ranking.
Evidence coverage
The auditable split is 10 described owner systems plus 3 independently attributable external sources. Repeated evidence from one physical machine counts once; separate machines can count separately even when they share a product model or owner. External packages remain labelled as sources when unique hardware identity cannot be safely proved.
Coverage includes Linux and Windows, thermals, power, multi-user serving, RPC, NPU sidecars, long context, vision and failed routes.
Claim-to-proof chain
Strengthen the map
Submit the exact system, firmware, runtime, command, model, quant, repeats and raw output. Community attribution and caveats stay attached.
These first-party direct llama-bench rows use locally built b10687 (c841aee),
kernel 7.0.0-30, Mesa/RADV 26.1.7 and the desktop performance profile. AMDGPU
DPM stayed on auto; the recorded CPU-only background workload remained active.
They are current-stack compatibility/speed observations, not replacements for
strict-clean headline runs or controlled comparisons against earlier kernels.
| Model | Quant | pp512 t/s | tg128 t/s | Repeats | Class |
|---|---|---|---|---|---|
| Qwen3-Coder 30B-A3B | UD-Q4_K_XL | 1264.16 | 94.64 | 20 | Sentinel |
| Qwen3-Next 80B-A3B | UD-Q4_K_XL | 675.76 | 62.09 | 20 | Sentinel |
| Qwen3.8-Flash-Next | UD-IQ4_XS | 394.73 | 27.16 | 10 | Single-artifact scout |
The Flash-Next GGUF is approximately 93.7GB. Its separate arithmetic smoke
returned 56 for 7×8 after the power profile was restored. Cite the repeated
benchmark for speed, not that smoke. Long context, vision, tools, server behavior
and broad task quality are not qualified. The
official model card counts
125B/6B active plus 51B n-gram embeddings and 4B MTP separately; the tested GGUF
reports about 177B parameters. It uses qwen-community-1.0, not Apache 2.0;
check the model terms before commercial deployment.
Evidence: structured rows, claim registry, sentinel notes and raw CSVs, Flash-Next notes and raw CSV, artifact hashes.
Strix Halo Guide is an independent community project. It is not affiliated with, endorsed by, or an official publication of AMD or any OEM. Product names identify relevant or tested hardware—not partnerships.
Read the disclosure policy