Series 01In progress / 2026

Local LLMs
on Strix Halo.

A product investigation from the hardware up: model fit, real throughput and the small decisions that decide whether a powerful local AI machine is genuinely useful.

ScopeEnd to end
PlatformUbuntu / Linux
RuntimesROCm / Vulkan / CPU
MeasureSpeed / power / fit

Why this machine

Memory changes the product conversation.

Strix Halo combines Zen 5 CPU cores, integrated RDNA 3.5 graphics and up to 128 GB of fast system memory. On the top configuration, AMD allows a large share of that memory to be assigned to graphics workloads.

That does not automatically make it a great LLM product platform. Software support, memory bandwidth, prompt processing, power and thermals still matter. This series is about the gap between can run, good to use and worth deploying.

Platform details are based on AMD’s published Ryzen AI Max+ 395 specifications. Test-machine details will be recorded in each result.

Publication roadmapUpdated as the work moves

Six notes, one reproducible path.

  1. 01Published

    GMKtec EVO-X3: the baseline

    A hardware inventory, provable firmware state and repeatable NVMe test from the theoretical link ceiling to a real model-file read.

  2. 02Planned

    ROCm without folklore

    A clean install, the working versions, the unsupported corners and a recovery path when the stack moves.

  3. 03Planned

    llama.cpp: Vulkan vs ROCm

    Prompt processing, generation speed, memory pressure and the flags that materially change behaviour.

  4. 04Planned

    Finding the useful quant

    Where model quality, context length and responsiveness meet for 30B, 70B and larger classes.

  5. 05Planned

    The long-session test

    Sustained generation, temperatures, clocks, power draw and whether the first result survives hour two.

  6. 06Planned

    A local stack worth keeping

    The final everyday setup: models, runtime, API layer, remote access, backups and honest limitations.

The test contract

Every result will ship with its context.