gputier

StrataGP — Development skill for Claude Code

Development community

Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux.

How to install StrataGP

This entry records only its repository, not the path inside it, so there is no exact command to give. Open gputier/StrataGP and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What StrataGP does

Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

Alternatives in Development

  • Magnitude — Open source inference server your agent sets up for you 1.5k ★
  • StatsPAI — StatsPAI is the first Agent-native Python library for causal inference and applied econometrics — unified API 302 ★
  • Claude Blocker — Block distracting websites unless Claude Code is actively running inference 282 ★

README

Strata

Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden that Strata's model wrote, running in the browser
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)

Strata runs **[Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next)** - a large, smart AI model that normally needs a server - on a normal PC. It chats, writes code, reads pictures and works with your apps and coding agents, and nothing leaves your PC.

How fast is it?

Measured on two ordinary gaming PCs. "Writes answers" is how fast the reply appears in a short chat; "reads your prompt" is how fast it takes in what you send (a 32K-token document, code or chat history). A token is about ¾ of a word, so 60 tokens per second is faster than you can read.

NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAMAMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
Size Writes answers Reads your prompt
Q2_0 94 tokens/s 2,650 tokens/s
IQ2_XS 79 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 1,620 tokens/s
Coder 55 tokens/s 2,180 tokens/s
Size Writes answers Reads your prompt
Q2_0 60 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 1,110 tokens/s
Coder 44 tokens/s 1,420 tokens/s

A card with more VRAM is faster: an RTX 3090 (24 GB) should write roughly 100-