0xSteph

LLM Rate Limit Proxy — AI skill for Claude Code

AI community

Self-hosted LLM proxy that pools and rotates your API keys, queues requests instead of returning 429 Too Many Requests, and fails over across providers.

How to install LLM Rate Limit Proxy

This entry records only its repository, not the path inside it, so there is no exact command to give. Open 0xSteph/llm-rate-limit-proxy and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What LLM Rate Limit Proxy does

Self-hosted LLM proxy that pools and rotates your API keys, queues requests instead of returning 429 Too Many Requests, and fails over across providers. Speaks both OpenAI and Anthropic, so Claude Code, Cline and Aider sit behind one endpoint. Single 4 MB Rust binary.

Alternatives in AI

  • Skill Seekers Roadmap — Transform Skill Seekers into the easiest way to create Claude AI skills from any knowledge source - documentat 11.1k ★
  • Simulink Agentic Toolkit — The Simulink Agentic Toolkit gives your AI agent both the tools and the expertise to work effectively with Sim 987 ★
  • Pretty Mermaid Skills — To provide AI with Mermaid chart rendering capability, supporting both SVG and ASCII output formats 552 ★

README

LLM Rate Limit Proxy

**Your coding agent never sees a rate limit again.**

Point Claude Code, Cline, Aider or any OpenAI-compatible harness at one endpoint with one key. It pools every API key you own across every provider you use, paces each request inside that key's real limit, and fails over when a key or a provider misbehaves — so the 429 that used to kill your agent mid-task never reaches it.

Speaks both the **OpenAI** (`/v1/chat/completions`) and **Anthropic** (`/v1/messages`) protocols, so Anthropic-native clients work without a shim.

Rust · single 4 MB binary · self-hosted · your keys never leave your machine.

[![CI](https://github.com/0xSteph/llm-rate-limit-proxy/actions/workflows/ci.yml/badge.svg)](https://github.com/0xSteph/llm-rate-limit-proxy/actions/workflows/ci.yml) [![Release](https://img.shields.io/github/v/release/0xSteph/llm-rate-limit-proxy?color=f0883e)](https://github.com/0xSteph/llm-rate-limit-proxy/releases) [![License](https://img.shields.io/badge/license-MIT%20OR%20Apache--2.0-blue)](#license)

![A 30-file refactor dying on a 429, then the same run completing through the proxy](docs/demo.gif)

Same agent. Same key. Same provider, same rate limit. The only difference is what sits in the middle.

The problem

Free and low-tier LLM APIs cap requests per minute, per key. Your agent burns through that cap in one refactor, the provider returns `429`, and the harness aborts — usually halfway through a multi-file edit, usually without saving its work.

The usual workarounds are all bad. Wait and retry by hand. Juggle three accounts and paste a different key each time. Pay for a tier you need for ten minutes a day.

Quickstart

curl -fsSL https://raw.githubusercontent.com/0xSteph/llm-rate-limit-proxy/master/install.sh | sh

Downloads the right binary, verifies its checksum, installs a hardened systemd service, and starts it. Open `http://localhost:8000`, and the wizard walks you through an admin account, your first provider key