Vllm Mlx — AI skill for Claude Code
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.
How to install Vllm Mlx
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open waybarrios/vllm-mlx and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Vllm Mlx does
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
Alternatives in AI
- Inference Serving — ai-research-skills vLLM, SGLang, TensorRT-LLM, llama.cpp 5.4k ★
- Deepreasoning — A high-performance LLM inference API and Chat UI that integrates DeepSeek R1's CoT reasoning traces with Anthr 5.4k ★
- Mac Code — mac code — Claude Code, but it runs on your Mac for free 905 ★
README
vllm-mlx
**Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference.**
**Read this in other languages:** [English](README.md) · [Español](README.es.md) · [Français](README.fr.md) · [中文](README.zh.md)
[](https://pypi.org/project/vllm-mlx/) [](https://pepy.tech/projects/vllm-mlx) [](https://www.python.org/downloads/) [](LICENSE) [](https://support.apple.com/en-us/HT211814) [](https://github.com/waybarrios/vllm-mlx)
What is vllm-mlx?
A vLLM-style inference server for Apple Silicon Macs. Unlike `Ollama` or `mlx-lm` used directly, it ships **continuous batching, paged KV cache, prefix caching, and SSD-tiered cache**, and exposes **both OpenAI `/v1/*` and Anthropic `/v1/messages`** from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step.
Quick start (30 seconds)
pip install vllm-mlx
vllm-mlx serve mlx-community/Llama-3.2-3B-Instruct-4bit --port 8000 --continuous-batching
**OpenAI SDK:**
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
r = client.chat.completions.create(model="default", messages=[{"role": "user", "content": "Hi!"}])
print(r.choices[0].message.content)
**Anthropic SDK / Claude Code:**
export ANTHROPIC_BASE_URL=http://localhost:8000
export ANTHROPIC_API_KEY=not-needed
claude
Features
APIs
- OpenAI-compatible:
/v1/chat/completions,/v1/completions,/v1/embeddings,/v1/rerank,/v1/responses - Anthropic-compatible: `/v1/mes
Related Skills
Yunshu
Fast local LLM / VLM inference engine for Apple Silicon (MLX). OpenAI- and Anthropic-compatible, lossless spec
Ovo Local LLM
A private Claude-Code-style coding agent for Apple Silicon — run chat, code, and local model workflows on-devi
Sous
Run Claude Code's subagents on a local LLM (MLX, Qwen) on Apple silicon and stretch your Pro/Max usage limits:
Setup Lmstudio VSCode
Set up a fully local coding model in VS Code on an Apple Silicon Mac (LM Studio MLX engine, no admin), with au
HomeAILab
Home AI inference lab running Qwen3.8 across RTX 5090, RTX 3090, and DGX Spark with vLLM, llama.cpp, custom Go
Wax
Shared Single-file memory layer for all your agents, sub mili-second RAG over text, photo and video on Apple S
Related Agents
LLM Integrator
LLM integration specialist who connects to OpenAI/Anthropic/Ollama APIs, designs prompt templates, implements
Fable5 Apfel Engineer
Heavy-lifting Fable-5 engineer for the apfel project (Apple on-device FoundationModels CLI + OpenAI-compatible
AI Recon
Delegates to this agent when the user wants to map the AI attack surface of an authorized web application befo