Octoroute โ AI skill for Claude Code
Smart HTTP router for local LLMs (Ollama, LM Studio, llama.cpp).
How to install Octoroute
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open slb350/octoroute and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Octoroute does
Smart HTTP router for local LLMs (Ollama, LM Studio, llama.cpp). Rule-based + LLM-powered routing, health checks, load balancing, Prometheus metrics. Rust-native, zero-overhead.
Alternatives in AI
- Repomix โ ๐ฆ Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k โ
- Ccstatusline โ by sirmalloc - A highly customizable status line formatter for Claude Code CLI that displays model info, git b 5.5k โ
- Inference Serving โ ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp 5.4k โ
README
Octoroute
Octoroute v3 is an OpenAI-compatible inference fabric. Clients use stable virtual model names while Octoroute selects an eligible local llama.cpp pool or an ordered provider target.
OpenAI-compatible client
|
v
Octoroute v3
/ \
local pools provider chain
The routing policy comes from configuration alone; Octoroute never classifies a prompt to choose a route. `X-Octoroute-Privacy: local-only` narrows a route before admission, so a local-only request cannot resolve provider credentials or disclose its prompt to a provider.
Current runtime
POST /v1/chat/completionswith schema-preserving request forwarding.GET /v1/models, liveness, readiness, and Prometheus endpoints.- Named virtual routes with ordered local-pool and provider steps.
- Exact local context/capability checks and least-loaded member selection.
- Per-member, per-provider, and inbound concurrency limits, plus an inbound per-minute request rate limit.
- Lazy, isolated provider credentials from environment variables or bounded argv commands.
- OpenAI-compatible HTTP dispatch for z.ai, OpenRouter, direct OpenAI, and similarly shaped endpoints.
- Explicit Anthropic Messages translation for text, tools, reasoning, non-streaming responses, and incremental SSE.
- Locked-down Codex CLI dispatch with ChatGPT-managed authentication, an allowlisted child environment, ephemeral read-only execution, and bounded structured output.
- An explicit OpenRouter Auto profile owned by Octoroute.
- Cached, bounded provider authentication/reachability probes and fixed-label Prometheus metrics for pool and provider admissions, responses, fallbacks, probes, and routing duration.
- Closed fallback triggers and a held first byte, preventing target changes after response commitment; an optional
first_byte_timeout_msbounds how long a hung upstream holds its permits before the route falls forward.
Quick start
Requirements:
- Rust 1.90 or ne
Related Skills
AIonDemandCluster
Spin up any open LLM on rented GPUs (vast.ai/RunPod), serve it with vLLM or llama.cpp, and drive it from Claud
Openclaw Self Healing
32+ 4-tier autonomous crash recovery for Claude Code and any service โ 64% auto-resolved, LLM-agnostic (Claude
Agent Observability Stack
Self-hostable observability for LLM agents + a Linux host (Intel iGPU/NPU + NVIDIA RTX 3090 eGPU aware): Prome
Axonhub
โก๏ธ Open-source AI Gateway โ Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & e
Llmio
LLM API load-balancing gateway. LLM API ่ด่ฝฝๅ่กก็ฝๅ ณ.
Lamoom Python
by LamoomAI - Serves as reference for production prompt engineering library with load balancing of AI Models,
Related Agents
Local Worker
Coding worker on a local model (Ollama, LM Studio, vLLM) served by the multi-model gateway's upstream. Use whe
Cpp Senior
[zakr] Senior C++ engineer. Use for C++ code review, RAII correctness, smart pointer usage, undefined behavior
Tron Integrator Sunswap
Use when integrating SunSwap DEX swaps on TRON โ building swap transactions via the Smart Exchange Router, enc