Director of ProductLondon, UK
I shape infrastructure products from the hardware up.
I’m Darren Soothill, a Director of Product working across infrastructure, hardware and storage. This is where product direction meets the bench: the systems I test, the trade-offs I find and the details that change the roadmap.
Current focus
Strix
Halo
Setup that can be repeated
Throughput in context
Thermals over time
Local AI on AMD Strix Halo
Capability · cost · experience · sustained performance
Product question / 01Local AI / 2026
When does local AI become a useful product choice?
01
Local LLMs on Strix Halo
The product question is bigger than a headline benchmark. Local inference changes privacy, cost, latency and deployment — but only if the hardware and software form a system people can actually use.
I took that question to the bench: what I installed, what broke, how I measured it and which engineering trade-offs deserve to shape the product decision.
Read the 9-part series-
01
GMKtec EVO-X3: the hardware, BIOS and NVMe baselinePublished
A product-led examination of my Linux-based GMKtec EVO-X3: Strix Halo hardware, its provable firmware state and measured NVMe performance.
-
02
ROCm on Strix Halo, without the folklorePublished
A measured Linux setup for ROCm on the Ryzen AI Max+ 395: memory, build provenance, model lifecycle, failure modes and a recovery runbook.
-
03
llama.cpp on Strix Halo: Vulkan versus ROCmPublished
Matched llama.cpp benchmarks show why Strix Halo has no single fastest backend: ROCm wins prompt processing while Vulkan can win generation.
-
04
Finding the useful quant: what fits is not what winsPublished
A measured 0.8B-to-397B Strix Halo model sweep showing why file size, active parameters, context cost and answer quality must be separate decisions.
-
05
The long-session test: 122B at a 256K contextPublished
A 101-minute, 165-request soak of Qwen3.5 122B at a nearly full 256K context, including throughput, memory, thermals and request-retention fixes.
-
06
A local LLM stack worth keepingPublished
The final Strix Halo stack: which hardware, Linux memory model, runtimes, models and operating controls I would keep after the benchmarks.
-
07
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deploymentPublished
A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads.
-
08
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching VulkanPublished
Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all four llama.cpp...
-
09
SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winnersPublished
Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.
The archive
Recent field notes
The technical detail behind product decisions: storage, Linux and the bits between a specification sheet and a working system.
View every noteSGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners
Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan
Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all...
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment
A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads....
A local LLM stack worth keeping
The final Strix Halo stack: which hardware, Linux memory model, runtimes, models and operating controls I would keep after the benchmarks.
How I workDirection ↔ detail
“Good product direction starts with the customer problem — and survives contact with the hardware.”
Start with the outcome
Who needs the product, what must change for them and which constraint is real.
Get close to the system
Hardware, storage paths, versions and topology reveal the actual trade space.
Bring evidence back
Measurements and rough edges turn assumptions into roadmap decisions.
Comparing notes on local AI or storage?
darren@soothill.com