IndexBy topic
Every field note, grouped by topic.
This index is generated from the article categories. For a guided route, use the dedicated Strix Halo or Storage page.
Automation
AWS
Benchmarking
Benchmarks
Qwen Vulkan PRs 28489 and 28501: below noise
Two draft changes delivered most of Qwen3.8's stack gain
Lemonade 11.9 on EVO-X3: the control plane changed, inference did not
Two llama.cpp Qwen fixes on Strix Halo: no free speed-up
Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production
Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan
Ornith 1.5 on Strix Halo: the 32K production profile
Qwen3.8 at 262K: fixing DFlash2 for production
Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost
Vulkan 0.6.4 on Strix Halo: the coding models moved
23 coding deployments on Strix Halo: what I would run
Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test
Strix Halo LLM upgrades: what got faster and what failed
SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment
The long-session test: 122B at a 256K context
Finding the useful quant: what fits is not what wins
llama.cpp on Strix Halo: Vulkan versus ROCm
Engineering
Qwen Vulkan PRs 28489 and 28501: below noise
Two draft changes delivered most of Qwen3.8's stack gain
Qwen3.8 can now roll back recurrent state, but that is only half the MTP story
The Qwen3.8 correctness merge I would carry before chasing speed
A Qwen3.8 prompt-cache slot can loop forever
Six Qwen3.8 MTP heads; only one fits my next test
Qwen3.8 MTP can mix users' answers across slots
Lemonade 11.9 on EVO-X3: the control plane changed, inference did not
Two llama.cpp Qwen fixes on Strix Halo: no free speed-up
Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production
Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan
Ornith 1.5 on Strix Halo: the 32K production profile
Qwen3.8 at 262K: fixing DFlash2 for production
Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost
Vulkan 0.6.4 on Strix Halo: the coding models moved
23 coding deployments on Strix Halo: what I would run
Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test
Strix Halo LLM upgrades: what got faster and what failed
SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment
Hardware
Linux
Local AI
Qwen Vulkan PRs 28489 and 28501: below noise
Two draft changes delivered most of Qwen3.8's stack gain
Qwen3.8 can now roll back recurrent state, but that is only half the MTP story
The Qwen3.8 correctness merge I would carry before chasing speed
Two QSA gather patches attack Qwen3.8's long-context decode cost
A Qwen3.8 prompt-cache slot can loop forever
Six Qwen3.8 MTP heads; only one fits my next test
Qwen3.8 MTP can mix users' answers across slots
A one-line Vulkan fix may unlock Q8 K/V prefill, but 125B still needs proving
Lazy auto mode can halve Qwen3.8 prefill on Strix Halo
Lemonade 11.9 on EVO-X3: the control plane changed, inference did not
Two llama.cpp Qwen fixes on Strix Halo: no free speed-up
Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production
Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan
Ornith 1.5 on Strix Halo: the 32K production profile
Qwen3.8 at 262K: fixing DFlash2 for production
Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost
Vulkan 0.6.4 on Strix Halo: the coding models moved
23 coding deployments on Strix Halo: what I would run
Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test
Strix Halo LLM upgrades: what got faster and what failed
SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners
DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan
EVO-X3 from BIOS to Lemonade: a reproducible build
DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment
GMKtec EVO-X3: the hardware, BIOS and NVMe baseline
A local LLM stack worth keeping
The long-session test: 122B at a 256K context
Finding the useful quant: what fits is not what wins
llama.cpp on Strix Halo: Vulkan versus ROCm
ROCm on Strix Halo, without the folklore
NVMe
Operations
Performance
Product
Reliability
S3
Software
SPDK
Storage
Upstream
Qwen3.8 can now roll back recurrent state, but that is only half the MTP story
The Qwen3.8 correctness merge I would carry before chasing speed
Two QSA gather patches attack Qwen3.8's long-context decode cost
A Qwen3.8 prompt-cache slot can loop forever
Six Qwen3.8 MTP heads; only one fits my next test
Qwen3.8 MTP can mix users' answers across slots
A one-line Vulkan fix may unlock Q8 K/V prefill, but 125B still needs proving
Lazy auto mode can halve Qwen3.8 prefill on Strix Halo