StrataGP — Development skill for Claude Code
Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux.
How to install StrataGP
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open gputier/StrataGP and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What StrataGP does
Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Alternatives in Development
- Magnitude — Open source inference server your agent sets up for you 1.5k ★
- StatsPAI — StatsPAI is the first Agent-native Python library for causal inference and applied econometrics — unified API 302 ★
- Claude Blocker — Block distracting websites unless Claude Code is actively running inference 282 ★
README
Strata
Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) ·
full video (49 s)
Strata runs **[Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next)** - a large, smart AI model that normally needs a server - on a normal PC. It chats, writes code, reads pictures and works with your apps and coding agents, and nothing leaves your PC.
How fast is it?
Measured on two ordinary gaming PCs. "Writes answers" is how fast the reply appears in a short chat; "reads your prompt" is how fast it takes in what you send (a 32K-token document, code or chat history). A token is about ¾ of a word, so 60 tokens per second is faster than you can read.
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
A card with more VRAM is faster: an RTX 3090 (24 GB) should write roughly 100-
Related Skills
Q Agent
Self-hosted llama.cpp stack for Qwen + MTP (multi-token prediction): a 125B MoE read from ≥ 80 GB of RAM, or a
Qwen3 Rs
A Qwen3.5 inference engine built by a group of agents on Slock.
Gvs5h
GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard Fable 5 Level Coding for a Fifth the
Nsight Graphics Analyzer
CLI wrapper and Claude Code / Codex skill for NVIDIA Nsight Graphics 2026.1+ — capture GPU frames, export GPU
GPU Check
Check GPU availability + driver version — ROCm untuk AMD RX 6700S, CUDA untuk NVIDIA GTX 1650.
Bio Doctor
Diagnostica a bio-stack — ferramentas, chaves, GPU, servidores MCP e endpoints NVIDIA — e diz exatamente o que
Related Agents
Nvidia Cuda Engineer
You are the system GPU Engineer — a senior NVIDIA platform specialist embedded in the system Distributed AI Ho
Lemonade Specialist
Lemonade Server and SDK specialist for local LLM on AMD hardware. Use PROACTIVELY for Lemonade setup, model ma
Aiml Engineer
AI/ML integration specialist. Designs model pipelines (inference, fine-tuning, LoRA adapters), selects infrast