Director of ProductLondon, UK

I shape infrastructure products from the hardware up.

I’m Darren Soothill, a Director of Product working across infrastructure, hardware and storage. This is where product direction meets the bench: the systems I test, the trade-offs I find and the details that change the roadmap.

Series / 01 Complete

Current focus

Strix
Halo

01

Setup that can be repeated

02

Throughput in context

03

Thermals over time

Current product lens

Local AI on AMD Strix Halo

Capability · cost · experience · sustained performance

Product question / 01Local AI / 2026

When does local AI become a useful product choice?

01

Local LLMs on Strix Halo

The product question is bigger than a headline benchmark. Local inference changes privacy, cost, latency and deployment — but only if the hardware and software form a system people can actually use.

I took that question to the bench: what I installed, what broke, how I measured it and which engineering trade-offs deserve to shape the product decision.

Read the 9-part series
  1. 01
    GMKtec EVO-X3: the hardware, BIOS and NVMe baseline

    A product-led examination of my Linux-based GMKtec EVO-X3: Strix Halo hardware, its provable firmware state and measured NVMe performance.

    Published
  2. 02
    ROCm on Strix Halo, without the folklore

    A measured Linux setup for ROCm on the Ryzen AI Max+ 395: memory, build provenance, model lifecycle, failure modes and a recovery runbook.

    Published
  3. 03
    llama.cpp on Strix Halo: Vulkan versus ROCm

    Matched llama.cpp benchmarks show why Strix Halo has no single fastest backend: ROCm wins prompt processing while Vulkan can win generation.

    Published
  4. 04
    Finding the useful quant: what fits is not what wins

    A measured 0.8B-to-397B Strix Halo model sweep showing why file size, active parameters, context cost and answer quality must be separate decisions.

    Published
  5. 05
    The long-session test: 122B at a 256K context

    A 101-minute, 165-request soak of Qwen3.5 122B at a nearly full 256K context, including throughput, memory, thermals and request-retention fixes.

    Published
  6. 06
    A local LLM stack worth keeping

    The final Strix Halo stack: which hardware, Linux memory model, runtimes, models and operating controls I would keep after the benchmarks.

    Published
  7. 07
    DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment

    A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads.

    Published
  8. 08
    DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan

    Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all four llama.cpp...

    Published
  9. 09
    SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners

    Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.

    Published

The archive

Recent field notes

The technical detail behind product decisions: storage, Linux and the bits between a specification sheet and a working system.

View every note
2026.08.10 local-ai

SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners

Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.

2026.08.09 local-ai

DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan

Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all...

2026.08.04 local-ai

DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment

A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads....

2026.08.03 local-ai

A local LLM stack worth keeping

The final Strix Halo stack: which hardware, Linux memory model, runtimes, models and operating controls I would keep after the benchmarks.

How I workDirection ↔ detail

“Good product direction starts with the customer problem — and survives contact with the hardware.”
01

Start with the outcome

Who needs the product, what must change for them and which constraint is real.

02

Get close to the system

Hardware, storage paths, versions and topology reveal the actual trade space.

03

Bring evidence back

Measurements and rough edges turn assumptions into roadmap decisions.

Comparing notes on local AI or storage?

darren@soothill.com