Local LLM Setup — AI skill for Claude Code
Runbook and PowerShell automation to host two isolated local LLMs (KYC analyst and KB-chatbot reasoning) on one 8 GB RTX GPU: Ollama on localhost behind a per-user bcrypt Basic-auth Caddy gateway, sin.
How to install Local LLM Setup
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open Ark310/local-llm-setup and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Local LLM Setup does
Runbook and PowerShell automation to host two isolated local LLMs (KYC analyst and KB-chatbot reasoning) on one 8 GB RTX GPU: Ollama on localhost behind a per-user bcrypt Basic-auth Caddy gateway, single-model VRAM policy, firewall and end-to-end tests.
Alternatives in AI
- Bot On Anything — A large model-based chatbot builder that can quickly integrate AI models (including ChatGPT, Claude, Gemini) i 4.2k ★
- Chatwise Releases — The fastest AI Chatbot for any LLM 1.4k ★
- Rb Setup — First-time setup 1.3k ★
README
Local LLM Setup: Two Isolated Models on One 8 GB GPU
Runbook and PowerShell automation for hosting two separate local LLM workloads (a KYC-analyst model and a knowledge-base chatbot's reasoning model) on one Windows gaming PC, with an authenticated, per-user LAN gateway in front of Ollama.
      
🧭 Part of **Abdul Raqeeb Khatri's portfolio**: [📂 Hub](https://github.com/Ark310/portfolio) · [🌐 Site](https://ark310.github.io) · [💼 Experience](https://github.com/Ark310/experience)
Overview
A team wanted a *local* reasoning model for its internal knowledge-base chatbot, served from a shared AI PC on the office LAN. The catch: the same PC (AMD Ryzen, **RTX 5060 Ti 8 GB**, 32 GB RAM) already hosts a **KYC-analyst model** whose prompts, test data and routes must never mix with the chatbot's. Two 7B-class models don't fit in 8 GB of VRAM together, and raw Ollama has no authentication.
This repo holds the rules, the design guide and the scripts that make the setup repeatable:
- One model server, two named models, one active at a time (
OLLAMA_MAX_LOADED_MODELS=1,ollama stopto switch). - Ollama stays on localhost. Colleagues reach the reasoning model only through a Caddy reverse proxy with per-user Basic auth.
- Hard separation between the KYC and chatbot folders, prompts, test data, logs and routes,
Related Skills
Empire AI Best Practices
Field-tested runbook + Claude Code skill for the Empire AI Alpha cluster: SSH/TOTP auth pattern for agents, Sl
Agent Observability Stack
Self-hostable observability for LLM agents + a Linux host (Intel iGPU/NPU + NVIDIA RTX 3090 eGPU aware): Prome
API VS Selfhost Skill
Anthropic-standard Skill — decide API-vs-self-host LLM costs and fine-tune ROI from any agent context (Claude
Skill Evolution
Self-improvement feedback loop for AI agent skills — analyzes past sessions, drafts structured improvement pro
Web Fetch
LLM-neutral skill that fetches a URL to a temp file so agents can pipe the path through rg / jq / awk instead
Claude Code Prompting 101
A comprehensive educational repository for mastering prompt engineering with Anthropic's Claude AI, from basic
Related Agents
Tech Debt Detector
Detecta deuda técnica en Booster AI — any, @ts-ignore, TODO/FIXME sin issue, localhost en código productivo, m
Local Worker
Coding worker on a local model (Ollama, LM Studio, vLLM) served by the multi-model gateway's upstream. Use whe
Compliance Director
コンプライアンスディレクターはカジノプラットフォームの規制対応、ライセンス要件、KYC/AML、責任あるギャンブル機能の最高責任者。全ての機能が規制要件を満たしていることを保証する。コンプライアンスに関わる決定や、規制リ