Yserver Local LLM System — AI skill for Claude Code
yserver: one small computer at home serves open LLMs to Claude Code and my coding tools.
How to install Yserver Local LLM System
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open YauhenBichel/yserver-local-llm-system and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Yserver Local LLM System does
yserver: one small computer at home serves open LLMs to Claude Code and my coding tools. The architecture, the flows, the measured numbers and the lessons. AMD Ryzen AI MAX+ 395 (Strix Halo), 128 GB shared memory, ROCm.
Alternatives in AI
- UI UX Pro Max Skill — An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms 45.8k ★
- Engineering Notebook — A Bun CLI that ingests Claude Code and Codex session transcripts, generates LLM-powered daily summaries, and s 320 ★
- Engram — The context spine that 10x's every AI coding session 141 ★
README
yserver: my local LLM system
One small computer at home, which I call yserver, serves open language models to my coding tools, my scripts and my side projects. This repository describes the system: what runs, how a request flows, how fast it is, and what went wrong.
It is a description, not an installer. The parts that are useful to other people are published as separate projects, and they are linked below.

**As a talk:** [One small computer, eight models](talk/) has the slides, a three-minute video and the script.
Why I built it
I use Claude Code, an editor assistant and my own scripts every day. I wanted them to work with models that I run myself. My prompts stay private, there is no usage limit, and there is no price per token. I still use the large cloud models for hard tasks. Much of my daily work does not need them.
The computer
| Part | What it is |
|---|---|
| Computer | GMKtec EVO-X2, a compact machine that I run as a server |
| Processor | AMD Ryzen AI MAX+ 395 ("Strix Halo"), with a built-in Radeon 8060S GPU |
| Memory | 128 GB. The CPU and the GPU share it. Half is reserved for the GPU, so a 52 GB model fits on a built-in GPU |
| System | Ubuntu 24.04, ROCm (the AMD software for GPU computing) |
What runs
| Layer | Part | What it does |
|---|---|---|
| Clients | Claude Code, an editor assistant, scripts, side projects | They speak the Anthropic API or the OpenAI API. They do not know which model answers |
| Entrance | A secure tunnel, then one gateway | The entrance for my tools. It checks access, chooses the model, checks the prompt length, and keeps a queue |
| Models | One model server on the GPU | One large model in memory at a time, with a context of 64,000 tokens |
| Models | Small servers for speech, transcription and images | Started when needed, stopped when idle |
| Agents | An agent runner | It gives a model a task, tools, a budget and checks. Every step i |
Related Skills
Sous
Run Claude Code's subagents on a local LLM (MLX, Qwen) on Apple silicon and stretch your Pro/Max usage limits:
1Helm
Self-hosted home for durable AI employees: one resident, one private computer, compounding memory and skills,
JakeyBot
GPT-5.3-powered multi-model Discord bot to interact with powerful LLMs and explore capabilities. This repo als
Agentic KVM
Give AI agents (Claude Code/Cowork, ChatGPT/Codex, Gemini/Antigravity) bare-metal, out-of-band control of any
Industry Watch
AI agent for semiconductor sector intelligence analysis. Tracks 8-K SEC filings and news for NVDA, AMD, INTC,
Vibe Halo
Windows approval popup and Dynamic Island for Codex, ZCode, Claude Code, OpenCode, and other AI coding agents.
Related Agents
Lesson Writer
Drafts the Hebrew lesson for a completed pipeline step from the step brief, report, diff and review — facts an
Peer Reviewer Computational
A simulated peer reviewer specializing in computational social science methods — NLP/text-as-data, machine lea
Search Issues
gh search issues --repo {owner/repo} "llms.txt OR AI-friendly OR LLM documentation OR AI context" --limit 20 g