IndexBy topic

Every field note, grouped by topic.

This index is generated from the article categories. For a guided route, use the dedicated Strix Halo or Storage page.

01

Automation

EVO-X3 from BIOS to Lemonade: a reproducible build
02

AWS

S3 benchmarking with parallel ranged downloads
03

Benchmarking

S3 benchmarking with parallel ranged downloads
04

Benchmarks

Qwen Vulkan PRs 28489 and 28501: below noise Two draft changes delivered most of Qwen3.8's stack gain Lemonade 11.9 on EVO-X3: the control plane changed, inference did not Two llama.cpp Qwen fixes on Strix Halo: no free speed-up Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan Ornith 1.5 on Strix Halo: the 32K production profile Qwen3.8 at 262K: fixing DFlash2 for production Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost Vulkan 0.6.4 on Strix Halo: the coding models moved 23 coding deployments on Strix Halo: what I would run Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test Strix Halo LLM upgrades: what got faster and what failed SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment The long-session test: 122B at a 256K context Finding the useful quant: what fits is not what wins llama.cpp on Strix Halo: Vulkan versus ROCm
05

Engineering

Qwen Vulkan PRs 28489 and 28501: below noise Two draft changes delivered most of Qwen3.8's stack gain Qwen3.8 can now roll back recurrent state, but that is only half the MTP story The Qwen3.8 correctness merge I would carry before chasing speed A Qwen3.8 prompt-cache slot can loop forever Six Qwen3.8 MTP heads; only one fits my next test Qwen3.8 MTP can mix users' answers across slots Lemonade 11.9 on EVO-X3: the control plane changed, inference did not Two llama.cpp Qwen fixes on Strix Halo: no free speed-up Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan Ornith 1.5 on Strix Halo: the 32K production profile Qwen3.8 at 262K: fixing DFlash2 for production Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost Vulkan 0.6.4 on Strix Halo: the coding models moved 23 coding deployments on Strix Halo: what I would run Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test Strix Halo LLM upgrades: what got faster and what failed SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment
06

Hardware

GMKtec EVO-X3: the hardware, BIOS and NVMe baseline
07

Linux

EVO-X3 from BIOS to Lemonade: a reproducible build ROCm on Strix Halo, without the folklore NVMe-oF with RoCE Configuration Guide - Ubuntu Server
08

Local AI

Qwen Vulkan PRs 28489 and 28501: below noise Two draft changes delivered most of Qwen3.8's stack gain Qwen3.8 can now roll back recurrent state, but that is only half the MTP story The Qwen3.8 correctness merge I would carry before chasing speed Two QSA gather patches attack Qwen3.8's long-context decode cost A Qwen3.8 prompt-cache slot can loop forever Six Qwen3.8 MTP heads; only one fits my next test Qwen3.8 MTP can mix users' answers across slots A one-line Vulkan fix may unlock Q8 K/V prefill, but 125B still needs proving Lazy auto mode can halve Qwen3.8 prefill on Strix Halo Lemonade 11.9 on EVO-X3: the control plane changed, inference did not Two llama.cpp Qwen fixes on Strix Halo: no free speed-up Qwen3.8 Flash Next: Vulkan 0.7.1 and Q8 in production Qwen3.8 Flash Next on AMD Strix Halo: ROCm vs Vulkan Ornith 1.5 on Strix Halo: the 32K production profile Qwen3.8 at 262K: fixing DFlash2 for production Vulkan 0.6.10 on Strix Halo: the MTP fix has a cost Vulkan 0.6.4 on Strix Halo: the coding models moved 23 coding deployments on Strix Halo: what I would run Muse Glimmer 30B on Strix Halo: Vulkan passed the 8K test Strix Halo LLM upgrades: what got faster and what failed SGLang vs vLLM vs llama.cpp on Strix Halo: five models, three winners DeepSeek V4 Flash on Strix Halo: tuning ROCm, matching Vulkan EVO-X3 from BIOS to Lemonade: a reproducible build DeepSeek V4 Flash 0731 on EVO-X3: the repeat that changed deployment GMKtec EVO-X3: the hardware, BIOS and NVMe baseline A local LLM stack worth keeping The long-session test: 122B at a 256K context Finding the useful quant: what fits is not what wins llama.cpp on Strix Halo: Vulkan versus ROCm ROCm on Strix Halo, without the folklore
09

NVMe

SPDK NVMe-oF target setup on Ubuntu with RDMA NVMe-oF with RoCE Configuration Guide - Ubuntu Server Changing an NVMe LBA format safely on Linux
10

Operations

A local LLM stack worth keeping
11

Performance

Two QSA gather patches attack Qwen3.8's long-context decode cost A one-line Vulkan fix may unlock Q8 K/V prefill, but 125B still needs proving Lazy auto mode can halve Qwen3.8 prefill on Strix Halo
12

Product

A local LLM stack worth keeping Finding the useful quant: what fits is not what wins
13

Reliability

The long-session test: 122B at a 256K context
14

S3

S3 benchmarking with parallel ranged downloads
15

Software

llama.cpp on Strix Halo: Vulkan versus ROCm ROCm on Strix Halo, without the folklore
16

SPDK

SPDK NVMe-oF target setup on Ubuntu with RDMA
17

Storage

GMKtec EVO-X3: the hardware, BIOS and NVMe baseline S3 benchmarking with parallel ranged downloads SPDK NVMe-oF target setup on Ubuntu with RDMA NVMe-oF with RoCE Configuration Guide - Ubuntu Server Changing an NVMe LBA format safely on Linux
18

Upstream

Qwen3.8 can now roll back recurrent state, but that is only half the MTP story The Qwen3.8 correctness merge I would carry before chasing speed Two QSA gather patches attack Qwen3.8's long-context decode cost A Qwen3.8 prompt-cache slot can loop forever Six Qwen3.8 MTP heads; only one fits my next test Qwen3.8 MTP can mix users' answers across slots A one-line Vulkan fix may unlock Q8 K/V prefill, but 125B still needs proving Lazy auto mode can halve Qwen3.8 prefill on Strix Halo