Local LLM Benchmarks — AI skill for Claude Code
Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB).
How to install Local LLM Benchmarks
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open kr4ckhe4d/local-llm-benchmarks and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Local LLM Benchmarks does
Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB). Qwen3.8, Gemma 4, GLM, gpt-oss; GGUF quant quality (KLD), MTP speculative decoding, 128K context, Claude Code with local models.
Alternatives in AI
- Repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k ★
- Qwen Code — A command-line AI workflow tool adapted from Gemini CLI, optimized for Qwen3-Coder models with enhanced parser 20.8k ★
- Inference Serving — ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp 5.4k ★
README
llama.cpp on AMD Radeon RDNA4 with ROCm: measured local LLM benchmarks
Real measurements of local LLMs on consumer AMD GPUs: an **RX 9070 XT (16 GB)** and a **Radeon AI PRO R9700 (32 GB)**, both RDNA4 / gfx1201, on llama.cpp with ROCm 7.2 under Linux. Throughput, VRAM fit, GGUF quantization quality (KL divergence against BF16), MTP speculative decoding, long context at real depth, and running coding agents (Claude Code, aider, Cline) against the local models. Every figure here was measured on this machine; scripts and raw logs are in [`benchmarks/`](benchmarks/).
**Hardware now:** Ryzen 7 9800X3D, 32 GB DDR5-6000, RX 9070 XT + Radeon AI PRO R9700 (48.9 GB VRAM, both PCIe 5.0 x8), CachyOS, ROCm 7.2, llama.cpp b11434.
Headline results (two cards, 2026-10)
| Model | Setup | Result |
|---|---|---|
| Gemma 4 26B-A4B | Q4 + MTP, 128K | 143 tok/s generation; 53 tok/s with 128K of context filled |
| Gemma 4 31B (dense) | Q6_K + MTP, 128K | 52 tok/s (19 without MTP); 279 tok/s prefill at 128K depth |
| Qwen3.6 35B-A3B | UD-Q6_K, 128K | 72 tok/s, only 2% below Q4 for 2.3x lower KLD; 3,170 tok/s prefill at 128K depth |
| Qwen3.8 27B (dense) | IQ4_XS + MTP | 71 tok/s, 85% draft acceptance (file since replaced by UD-Q6_K) |
| Qwen3.8 27B | UD-Q6_K + MTP | 57 tok/s; KLD vs BF16 0.0020 (IQ4_XS: 0.0179); as good as Q8_0 in Claude Code, 12% faster |
| GLM-4.7-Flash 30B-A3B | Q4/Q8, 128K | 75 tok/s in an empty context, but 154 tok/s prefill at 128K depth (attention-bound; removed) |
| MoE models generally | second card | CPU expert offload gone: +38% to +183% generation |
Generation is the 700-token probe in an empty context unless it says "depth"; filled to 128K, every model keeps 20-51% of that rate. Details and caveats are in the linked files.
Where to look
| File | What it answers |
|---|---|
| r9700+rx9070.md | What a second GPU (R9700 beside a 9070 XT) changed: every preset re-fit, MoE without CPU offload, dense models w |
Related Skills
HomeAILab
Home AI inference lab running Qwen3.8 across RTX 5090, RTX 3090, and DGX Spark with vLLM, llama.cpp, custom Go
Yserver Local LLM System
yserver: one small computer at home serves open LLMs to Claude Code and my coding tools. The architecture, the
Roughcut AI Local First Editor
Open-source, local-first AI video editor: on-device transcription (whisper.cpp), local LLM rough cuts (Gemma v
OmniGlyph
Cut your Claude bill 59–70% by rendering bulky LLM context as dense PNG pages — 100% read accuracy, exact per-
AIonDemandCluster
Spin up any open LLM on rented GPUs (vast.ai/RunPod), serve it with vLLM or llama.cpp, and drive it from Claud
Quackd
🦆🧠 Give your Microduck a brain. Tell a small robot with two legs what you want in plain language. An LLM (Cl