LLM Rate Limit Proxy — AI skill for Claude Code
Self-hosted LLM proxy that pools and rotates your API keys, queues requests instead of returning 429 Too Many Requests, and fails over across providers.
How to install LLM Rate Limit Proxy
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open 0xSteph/llm-rate-limit-proxy and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What LLM Rate Limit Proxy does
Self-hosted LLM proxy that pools and rotates your API keys, queues requests instead of returning 429 Too Many Requests, and fails over across providers. Speaks both OpenAI and Anthropic, so Claude Code, Cline and Aider sit behind one endpoint. Single 4 MB Rust binary.
Alternatives in AI
- Skill Seekers Roadmap — Transform Skill Seekers into the easiest way to create Claude AI skills from any knowledge source - documentat 11.1k ★
- Simulink Agentic Toolkit — The Simulink Agentic Toolkit gives your AI agent both the tools and the expertise to work effectively with Sim 987 ★
- Pretty Mermaid Skills — To provide AI with Mermaid chart rendering capability, supporting both SVG and ASCII output formats 552 ★
README
LLM Rate Limit Proxy
**Your coding agent never sees a rate limit again.**
Point Claude Code, Cline, Aider or any OpenAI-compatible harness at one endpoint with one key. It pools every API key you own across every provider you use, paces each request inside that key's real limit, and fails over when a key or a provider misbehaves — so the 429 that used to kill your agent mid-task never reaches it.
Speaks both the **OpenAI** (`/v1/chat/completions`) and **Anthropic** (`/v1/messages`) protocols, so Anthropic-native clients work without a shim.
Rust · single 4 MB binary · self-hosted · your keys never leave your machine.
[](https://github.com/0xSteph/llm-rate-limit-proxy/actions/workflows/ci.yml) [](https://github.com/0xSteph/llm-rate-limit-proxy/releases) [](#license)

Same agent. Same key. Same provider, same rate limit. The only difference is what sits in the middle.
The problem
Free and low-tier LLM APIs cap requests per minute, per key. Your agent burns through that cap in one refactor, the provider returns `429`, and the harness aborts — usually halfway through a multi-file edit, usually without saving its work.
The usual workarounds are all bad. Wait and retry by hand. Juggle three accounts and paste a different key each time. Pay for a tier you need for ten minutes a day.
Quickstart
curl -fsSL https://raw.githubusercontent.com/0xSteph/llm-rate-limit-proxy/master/install.sh | sh
Downloads the right binary, verifies its checksum, installs a hardened systemd service, and starts it. Open `http://localhost:8000`, and the wizard walks you through an admin account, your first provider key
Related Skills
Agent Queue
Task queue and orchestrator for AI coding agents. Manage Claude Code agents from Discord — auto-recovers from
Folder Lock
Many AI agents, one repo, no conflicting saves - a Claude Code skill: one lock per workfolder, handoffs instea
Headroom AI
macOS menu bar app showing rate limits across many Claude Code and Codex accounts at once
Task Automation AI Agent
A Python agent that breaks natural-language requests into sub-tasks and calls tools (calculator, knowledge sea
Echook
🔊 echook — AI-operated audio notifications for Claude Code, Cursor IDE & Codex CLI — 26 hooks, voice + chime
Tmux Agent Usage
Display AI agent rate limit usage in your tmux status bar
Related Agents
ZeroClaw Android
Run AI agents 24/7 on your Android phone. Native Rust core, 25+ providers (OpenAI, Claude, Gemini, Groq, DeepS
Castari Proxy
Use Claude Agent SDK and Claude Code with other providers/models.
Codex Debugger
Root-cause a failing test, crash, stack trace, or misbehaving feature by delegating the investigation to OpenA