Gemini Model Router — AI skill for Claude Code
Local-first agentic terminal that routes prompts across local Gemma 4 (vLLM), Gemini CLI, and Claude Code — picking the right one per cost/latency/quality.
How to install Gemini Model Router
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open jswortz/gemini-model-router and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Gemini Model Router does
Local-first agentic terminal that routes prompts across local Gemma 4 (vLLM), Gemini CLI, and Claude Code — picking the right one per cost/latency/quality. Logs every decision in gemini-dreams shape.
Alternatives in AI
- Repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k ★
- Inference Serving — ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp 5.4k ★
- Vllm Mlx — High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon 1.5k ★
README
gemini-model-router
A local-first agentic terminal that routes each prompt across **local Gemma 4 (vLLM)**, **Gemini CLI**, and **Claude Code** — picking the right one to optimize cost, latency, and quality. Logs every decision in a format the [gemini-dreams](https://github.com/jswortz/gemini-dreams) nightly loop can mine to propose its own improvements.
[](https://asciinema.org/a/P3pkHoXam0N1MJKM)
Install
Recommended — install as a uv tool so the scripts land on `PATH` and work from any directory in any new shell:
cd gemini-model-router
uv tool install -e .
uv tool update-shell # one-time PATH fix; open a new terminal after
Alternative — editable install into a project venv (you must activate the venv or add `.venv/bin` to `PATH` to use `router`):
uv pip install -e ".[dev]"
source .venv/bin/activate
This wires four console scripts: `router`, `router-repl`, `router-eval`, `router-config`. See [`docs/cli.md`](docs/cli.md#install) for extras and upgrade flow.
Prerequisites
- Local vLLM serving Gemma 4 on
localhost:8000:vllm serve google/gemma-4-E4B-it --port 8000 geminiCLI onPATH, withGEMINI_API_KEY(or Google Sign-in) configured.claudeCLI onPATH, withANTHROPIC_API_KEYconfigured.- (Optional)
HF_TOKENfor HuggingFace — silences the unauthenticated-Hub warning when the MiniLM anchor cache is rebuilt.
Override any of these in `config/router.yaml`. Or drop them in a `.env` (see `.env.example`); the router auto-loads it on startup, never overriding already-exported shell vars.
Usage
**One-shot mode:**
router "what is the capital of France"
router "refactor src/auth/login.py to use async/await" --why
router --force gemini "search the web for the latest k8s release notes"
**Chat mode (TUI)** — `router` with no prompt at a TTY drops into a chat REPL with slash-command completion, a live status t
Related Skills
Claude Code Jev Smart Router
HTTP proxy for Claude Code that selects the Claude model per request to cut cost and latency. Routes on task p
Laya Fast Router
A lightweight decision routing model designed to replace slow, expensive LLMs for simple classification and co
Token Ninja
token-ninja routes deterministic shell commands locally — zero LLM calls, ~19µs latency. Works silently inside
Local LLM Benchmarks
Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB). Qwen3.8
AIonDemandCluster
Spin up any open LLM on rented GPUs (vast.ai/RunPod), serve it with vLLM or llama.cpp, and drive it from Claud
Aurora
Aurora - independent community smart home orchestrator. Routes any smart home request to the right specialist
Related Agents
Staff Software Engineer
Engineering orchestrator. Runs the full pipeline (plan, dev, test, pr) for a feature, picking the right area s
Brain Eval Engineer
Evaluation engineer — question set tooling, layered metrics (harvest/graph/retrieval/answer), fixed-strategy a
Forge Canary
Forge's post-deploy monitoring agent — watches a deployed URL for a time window right after release, tracking