Darren SoothillNewport Pagnell, UK

I test emerging technology to find out what it can actually do.

I’m a Director of Product working with storage and GPU-based systems. This site is my record of the questions I asked, the tests I ran, what failed and the point at which the result became useful for a product decision.

Series / 01 Ongoing

Current focus

Strix
Halo

01

Compatibility

02

Sustained behaviour

03

Product decision

Current investigationOngoing / 2026

Strix Halo can run large local models. That does not settle whether the stack is mature.

Strix Halo

Local LLMs on Strix Halo

Runtime support is changing quickly as compatibility expands. A model loading once tells me very little about correctness, memory behaviour, long sessions or recovery after a failure.

I record the hardware, build, launch settings and workload, then keep compatibility, performance, maturity and stability as separate questions. The conclusion stays attached to the conditions that produced it.

Read the ongoing investigation and 10 articles published so far
  1. 01
    23 coding deployments on Strix Halo: what I would run

    I ran 23 local coding-model deployments through 230 executable tasks on a 128GB Strix Halo workstation, comparing accuracy, latency and context limits.

    Newest
  2. 02
    SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners

    Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.

    Recent
  3. 03
    DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan

    Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all four llama.cpp...

    Recent
Show 7 more articles Hide the older articles
  1. 04
    DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment

    A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads.

    Earlier
  2. 05
    GMKtec EVO-X3: the hardware, BIOS and NVMe baseline

    A product-led examination of my Linux-based GMKtec EVO-X3: Strix Halo hardware, its provable firmware state and measured NVMe performance.

    Earlier
  3. 06
    A local LLM stack worth keeping

    The Strix Halo stack I would keep after the first benchmark sequence: hardware, Linux memory, runtimes, models and operating controls.

    Earlier
  4. 07
    The long-session test: 122B at a 256K context

    A 101-minute, 165-request soak of Qwen3.5 122B at a nearly full 256K context, including throughput, memory, thermals and request-retention fixes.

    Earlier
  5. 08
    Finding the useful quant: what fits is not what wins

    A measured 0.8B-to-397B Strix Halo model sweep showing why file size, active parameters, context cost and answer quality must be separate decisions.

    Earlier
  6. 09
    llama.cpp on Strix Halo: Vulkan versus ROCm

    Matched llama.cpp benchmarks show why Strix Halo has no single fastest backend: ROCm wins prompt processing while Vulkan can win generation.

    Earlier
  7. 10
    ROCm on Strix Halo, without the folklore

    A measured Linux setup for ROCm on the Ryzen AI Max+ 395: memory, build provenance, model lifecycle, failure modes and a recovery runbook.

    Earlier

What changedSelected results

The useful answers came from looking past the first successful run.

23 deployments / 230 requests

Completion was not the decision

Every scored coding request completed, but the deployments still differed in correctness, speed and practical fit. The recommendation had to account for all three.

Read the comparison
4.09 hours

The longer run narrowed the claim

One sustained DeepSeek run showed no progressive leak signal within that test. It improved confidence in that configuration; it did not prove indefinite stability.

Read the test record

Storage and infrastructure

The same questions apply below the AI stack.

The storage work covers the details that decide whether a system is repeatable and serviceable: device identity, access control, network assumptions, recovery and the exact workload behind a performance number.

Explore the storage work
GuideSPDK

SPDK NVMe-oF target with RDMA

A controlled lab path with a pinned release, disposable storage and explicit host access.

GuideNVMe-oF

Kernel NVMe-oF over RoCE

A persistent target built around stable device identities, visible data-loss boundaries and end-to-end network checks.

ToolS3

Parallel ranged-download benchmarking

A Go tool for testing where the client, network and object-store path reach their limit.

How I work

I start with the claim, test the mechanism underneath it and state where the result stops.

That is the link between my product work and the technical material here. A benchmark is useful when it changes a decision and its limits are clear. Read more about the approach

Working through a storage, GPU or local-AI product question?

darren@soothill.com