Inference Architect — DevOps & Infrastructure agent for Claude Code
You are the system Inference Architect — a senior ML infrastructure engineer specializing in local LLM serving, model optimization, GPU resource management, and inference routing across heterogene.
How to install Inference Architect
Installs to ~/.claude/agents/viriansemail-sys-claude-code-helper-inference-architect.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/viriansemail-sys/Claude-Code-Helper/HEAD/agents/inference-architect.md -o ~/.claude/agents/viriansemail-sys-claude-code-helper-inference-architect.md Restart Claude Code, or start a new session, for it to be picked up.
What Inference Architect does
name: inference-architect model: opus color: red memory: user
You are **the system Inference Architect** — a senior ML infrastructure engineer specializing in local LLM serving, model optimization, GPU resource management, and inference routing across heterogeneous NVIDIA compute clusters. You understand that the system's intelligence runs on *local iron* — not cloud APIs — and every architectural decision you make affects response latency, model quality, and power consumption.
You thi
Alternatives in DevOps & Infrastructure
- Kubernetes Architect — You are a Kubernetes architect specializing in cloud-native infrastructure, modern GitOps workflows 31.9k ★
- Claudable — by Ethan Park - Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code a 3.8k ★
- Cloud Architect — Cloud architect for AWS/Azure/GCP infrastructure, IaC, FinOps, and multi-cloud strategies 2.8k ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Aiml Engineer
AI/ML integration specialist. Designs model pipelines (inference, fine-tuning, LoRA adapters), selects infrast
MLOps
ML infrastructure specialist — training pipelines, model serving, GPU optimization, distributed training
CI Architect
CI/CD pipeline architect specializing in workflow configuration, job orchestration, and pipeline optimization.
MLOps Reviewer
MLOps and ML-infrastructure expert for the code-review panel. Spawned when the diff touches pipelines (Airflow
AWS Integration
Configure and manage AWS integration for monitoring, log collection, and resource tracking across AWS accounts
Azure Integration
Configure and manage Azure integration for monitoring, log collection, and resource tracking across Azure subs
Related Skills
Samzong
Inference Serving Stack: Scheduling and routing across heterogeneous models — the production path for vLLM and
Token Optimizer Skill
Token cost optimization skill for Claude Code — model routing (Opus/Sonnet/Haiku), extended thinking tuning, p
Inference Serving
ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp