YauhenBichel

Yserver Local LLM System — AI skill for Claude Code

AI community

yserver: one small computer at home serves open LLMs to Claude Code and my coding tools.

How to install Yserver Local LLM System

This entry records only its repository, not the path inside it, so there is no exact command to give. Open YauhenBichel/yserver-local-llm-system and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Yserver Local LLM System does

yserver: one small computer at home serves open LLMs to Claude Code and my coding tools. The architecture, the flows, the measured numbers and the lessons. AMD Ryzen AI MAX+ 395 (Strix Halo), 128 GB shared memory, ROCm.

Alternatives in AI

  • UI UX Pro Max Skill — An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms 45.8k ★
  • Engineering Notebook — A Bun CLI that ingests Claude Code and Codex session transcripts, generates LLM-powered daily summaries, and s 320 ★
  • Engram — The context spine that 10x's every AI coding session 141 ★

README

yserver: my local LLM system

One small computer at home, which I call yserver, serves open language models to my coding tools, my scripts and my side projects. This repository describes the system: what runs, how a request flows, how fast it is, and what went wrong.

It is a description, not an installer. The parts that are useful to other people are published as separate projects, and they are linked below.

![Architecture of my own LLM system](diagrams/architecture.png)

**As a talk:** [One small computer, eight models](talk/) has the slides, a three-minute video and the script.

Why I built it

I use Claude Code, an editor assistant and my own scripts every day. I wanted them to work with models that I run myself. My prompts stay private, there is no usage limit, and there is no price per token. I still use the large cloud models for hard tasks. Much of my daily work does not need them.

The computer

Part What it is
Computer GMKtec EVO-X2, a compact machine that I run as a server
Processor AMD Ryzen AI MAX+ 395 ("Strix Halo"), with a built-in Radeon 8060S GPU
Memory 128 GB. The CPU and the GPU share it. Half is reserved for the GPU, so a 52 GB model fits on a built-in GPU
System Ubuntu 24.04, ROCm (the AMD software for GPU computing)

What runs

Layer Part What it does
Clients Claude Code, an editor assistant, scripts, side projects They speak the Anthropic API or the OpenAI API. They do not know which model answers
Entrance A secure tunnel, then one gateway The entrance for my tools. It checks access, chooses the model, checks the prompt length, and keeps a queue
Models One model server on the GPU One large model in memory at a time, with a context of 64,000 tokens
Models Small servers for speech, transcription and images Started when needed, stopped when idle
Agents An agent runner It gives a model a task, tools, a budget and checks. Every step i