kr4ckhe4d

Local LLM Benchmarks — AI skill for Claude Code

AI community

Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB).

How to install Local LLM Benchmarks

This entry records only its repository, not the path inside it, so there is no exact command to give. Open kr4ckhe4d/local-llm-benchmarks and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Local LLM Benchmarks does

Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB). Qwen3.8, Gemma 4, GLM, gpt-oss; GGUF quant quality (KLD), MTP speculative decoding, 128K context, Claude Code with local models.

Alternatives in AI

  • Repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k ★
  • Qwen Code — A command-line AI workflow tool adapted from Gemini CLI, optimized for Qwen3-Coder models with enhanced parser 20.8k ★
  • Inference Serving — ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp 5.4k ★

README

llama.cpp on AMD Radeon RDNA4 with ROCm: measured local LLM benchmarks

Real measurements of local LLMs on consumer AMD GPUs: an **RX 9070 XT (16 GB)** and a **Radeon AI PRO R9700 (32 GB)**, both RDNA4 / gfx1201, on llama.cpp with ROCm 7.2 under Linux. Throughput, VRAM fit, GGUF quantization quality (KL divergence against BF16), MTP speculative decoding, long context at real depth, and running coding agents (Claude Code, aider, Cline) against the local models. Every figure here was measured on this machine; scripts and raw logs are in [`benchmarks/`](benchmarks/).

**Hardware now:** Ryzen 7 9800X3D, 32 GB DDR5-6000, RX 9070 XT + Radeon AI PRO R9700 (48.9 GB VRAM, both PCIe 5.0 x8), CachyOS, ROCm 7.2, llama.cpp b11434.

Headline results (two cards, 2026-10)

Model Setup Result
Gemma 4 26B-A4B Q4 + MTP, 128K 143 tok/s generation; 53 tok/s with 128K of context filled
Gemma 4 31B (dense) Q6_K + MTP, 128K 52 tok/s (19 without MTP); 279 tok/s prefill at 128K depth
Qwen3.6 35B-A3B UD-Q6_K, 128K 72 tok/s, only 2% below Q4 for 2.3x lower KLD; 3,170 tok/s prefill at 128K depth
Qwen3.8 27B (dense) IQ4_XS + MTP 71 tok/s, 85% draft acceptance (file since replaced by UD-Q6_K)
Qwen3.8 27B UD-Q6_K + MTP 57 tok/s; KLD vs BF16 0.0020 (IQ4_XS: 0.0179); as good as Q8_0 in Claude Code, 12% faster
GLM-4.7-Flash 30B-A3B Q4/Q8, 128K 75 tok/s in an empty context, but 154 tok/s prefill at 128K depth (attention-bound; removed)
MoE models generally second card CPU expert offload gone: +38% to +183% generation

Generation is the 700-token probe in an empty context unless it says "depth"; filled to 128K, every model keeps 20-51% of that rate. Details and caveats are in the linked files.

Where to look

File What it answers
r9700+rx9070.md What a second GPU (R9700 beside a 9070 XT) changed: every preset re-fit, MoE without CPU offload, dense models w