Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Sep 8, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
The fastest way to run Qwen3.8 27B on Strix Halo (gfx1151)
Local text to textured GLB on an AMD Strix Halo iGPU (gfx1151): FLUX.2 klein, a Vulkan-only TRELLIS.2 engine, and humanoid auto-rigging with SkinTokens on ROCm. No Blender, no CUDA.
Docker Compose for llama.cpp GGUF servers on AMD Strix Halo: Qwen, Gemma, and Laguna packages (abliterated and quantized), stock Vulkan plus ROCmFP4/MTP and ROCmFPX, parallel slots, with prefill/decode and quality metrics measured on this rig.
Tensor-parallel DeepSeek V4 Flash inference on dual AMD Strix Halo over OdinLink USB4/TB5 RDMA or Mellanox RoCE v2.
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
Clean-room AGPL-3.0 reimplementation of the halogen 0.1.3 inference engine for Qwen3.8-27B on AMD Strix Halo (gfx1151) — same wire protocol and OpenAI-compatible API. Not affiliated with Peonist.
llama.cpp + Qwen3.6-27B (Q8_0 GGUF) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). 256K context, ~7.5 t/s decode via TheRock ROCm Docker.
Claude Code skill for AMD Strix Halo (Ryzen AI MAX+ 395) ML setup. Handles PyTorch installation (official wheels don't work with gfx1151), GTT memory config, and environment setup. Enables 30B parameter models.
Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
Measured LLM benchmarks for AMD Strix Halo / Ryzen AI Max+ 395 (Radeon 8060S, 128 GB unified): llama.cpp Vulkan & ROCm — decode pace, TTFA, prompt cache, quants, sustained load. Every number links to raw runs.
Run large LLMs locally on AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151) with ROCmFP4 4-bit quantization. Measured benchmarks, build + serving recipes, and 118 ready-to-run GGUF models.
To associate your repository with the gfx1151 topic, visit your repo's landing page and select "manage topics."