Director of ProductLondon, UK
I shape infrastructure products from the hardware up.
I’m Darren Soothill, a Director of Product working across infrastructure, hardware and storage. This is where product direction meets the bench: the systems I test, the trade-offs I find and the details that change the roadmap.
Current focus
Strix
Halo
Setup that can be repeated
Throughput in context
Thermals over time
Local AI on AMD Strix Halo
Capability · cost · experience · sustained performance
Product question / 01Local AI / 2026
When does local AI become a useful product choice?
01
Local LLMs on Strix Halo
The product question is bigger than a headline benchmark. Local inference changes privacy, cost, latency and deployment — but only if the hardware and software form a system people can actually use.
I took that question to the bench: what I installed, what broke, how I measured it and which engineering trade-offs deserve to shape the product decision.
Read the 7-part series-
01
GMKtec EVO-X3 baselinePublished
Hardware, firmware state and NVMe performance measured as one Linux system.
-
02
ROCm without folklorePublished
The working Linux stack, its support boundary and a recovery runbook.
-
03
Vulkan versus ROCmPublished
Matched prompt, generation and concurrency results across two models.
-
04
Find the useful quantPublished
Fit, context and speed from 0.8B to a 397B-class model.
-
05
Run the long-session testPublished
122B at a near-full 256K context for more than 100 minutes.
-
06
Keep the useful stackPublished
The final product, runtime and operations decisions.
-
07
DeepSeek V4 Flash 0731 on evox3Published
A failed first load, a measured 10× experiment and the repeat test that changed deployment.
The archive
Recent field notes
The technical detail behind product decisions: storage, Linux and the bits between a specification sheet and a working system.
View every noteDeepSeek V4 Flash 0731 on evox3: the repeat that changed deployment
DeepSeek V4 Flash 0731 on evox3 and ROCm 7.14: a failed first load, a measured 10× sparse-prefill result, and the repeat test...
A local LLM stack worth keeping
The final Strix Halo stack: which hardware, Linux memory model, runtimes, models and operating controls I would keep after the benchmarks.
The long-session test: 122B at a 256K context
A 101-minute, 165-request soak of Qwen3.5 122B at a nearly full 256K context, including throughput, memory, thermals and request-retention fixes.
Finding the useful quant: what fits is not what wins
A measured 0.8B-to-397B Strix Halo model sweep showing why file size, active parameters, context cost and answer quality must be separate decisions....
How I workDirection ↔ detail
“Good product direction starts with the customer problem — and survives contact with the hardware.”
Start with the outcome
Who needs the product, what must change for them and which constraint is real.
Get close to the system
Hardware, storage paths, versions and topology reveal the actual trade space.
Bring evidence back
Measurements and rough edges turn assumptions into roadmap decisions.
Comparing notes on local AI or storage?
darren@soothill.com