Darren SoothillNewport Pagnell, UK
I test emerging technology to find out what it can actually do.
I’m a Director of Product working with storage and GPU-based systems. This site is my record of the questions I asked, the tests I ran, what failed and the point at which the result became useful for a product decision.
Current focus
Strix
Halo
Compatibility
Sustained behaviour
Product decision
Current investigationOngoing / 2026
Strix Halo can run large local models. That does not settle whether the stack is mature.
Strix Halo
Local LLMs on Strix Halo
Runtime support is changing quickly as compatibility expands. A model loading once tells me very little about correctness, memory behaviour, long sessions or recovery after a failure.
I record the hardware, build, launch settings and workload, then keep compatibility, performance, maturity and stability as separate questions. The conclusion stays attached to the conditions that produced it.
Read the ongoing investigation and 10 articles published so far-
01
23 coding deployments on Strix Halo: what I would runNewest
I ran 23 local coding-model deployments through 230 executable tasks on a 128GB Strix Halo workstation, comparing accuracy, latency and context limits.
-
02
SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winnersRecent
Twenty matched SGLang-vLLM points plus llama.cpp and DeepSeek tests show why vLLM wins native Qwen while specialised runtimes win large quants.
-
03
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching VulkanRecent
Matched DeepSeek V4 Flash tests on Strix Halo show a 44% ROCm prefill gain from tuning, but patched Vulkan still wins all four llama.cpp...
Show 7 more articles Hide the older articles
-
04
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deploymentEarlier
A measured DeepSeek V4 Flash 0731 study on Strix Halo: ROCm optimisation, a four-hour 32K thermal soak and 20 cold model loads.
-
05
GMKtec EVO-X3: the hardware, BIOS and NVMe baselineEarlier
A product-led examination of my Linux-based GMKtec EVO-X3: Strix Halo hardware, its provable firmware state and measured NVMe performance.
-
06
A local LLM stack worth keepingEarlier
The Strix Halo stack I would keep after the first benchmark sequence: hardware, Linux memory, runtimes, models and operating controls.
-
07
The long-session test: 122B at a 256K contextEarlier
A 101-minute, 165-request soak of Qwen3.5 122B at a nearly full 256K context, including throughput, memory, thermals and request-retention fixes.
-
08
Finding the useful quant: what fits is not what winsEarlier
A measured 0.8B-to-397B Strix Halo model sweep showing why file size, active parameters, context cost and answer quality must be separate decisions.
-
09
llama.cpp on Strix Halo: Vulkan versus ROCmEarlier
Matched llama.cpp benchmarks show why Strix Halo has no single fastest backend: ROCm wins prompt processing while Vulkan can win generation.
-
10
ROCm on Strix Halo, without the folkloreEarlier
A measured Linux setup for ROCm on the Ryzen AI Max+ 395: memory, build provenance, model lifecycle, failure modes and a recovery runbook.
What changedSelected results
The useful answers came from looking past the first successful run.
Completion was not the decision
Every scored coding request completed, but the deployments still differed in correctness, speed and practical fit. The recommendation had to account for all three.
Read the comparisonThe longer run narrowed the claim
One sustained DeepSeek run showed no progressive leak signal within that test. It improved confidence in that configuration; it did not prove indefinite stability.
Read the test recordStorage and infrastructure
The same questions apply below the AI stack.
The storage work covers the details that decide whether a system is repeatable and serviceable: device identity, access control, network assumptions, recovery and the exact workload behind a performance number.
Explore the storage workSPDK NVMe-oF target with RDMA
A controlled lab path with a pinned release, disposable storage and explicit host access.
Kernel NVMe-oF over RoCE
A persistent target built around stable device identities, visible data-loss boundaries and end-to-end network checks.
Parallel ranged-download benchmarking
A Go tool for testing where the client, network and object-store path reach their limit.
How I work
I start with the claim, test the mechanism underneath it and state where the result stops.
That is the link between my product work and the technical material here. A benchmark is useful when it changes a decision and its limits are clear. Read more about the approach
Working through a storage, GPU or local-AI product question?
darren@soothill.com