Mac Code — AI skill for Claude Code
mac code — Claude Code, but it runs on your Mac for free.
How to install Mac Code
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open walter-grace/mac-code and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Mac Code does
mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.
Alternatives in AI
- System Prompts Leaks — Extracted system prompts from ChatGPT (GPT-5.4, GPT-5.3, Codex), Claude (Opus 4.6, Sonnet 4.6, Claude Code), G 38.6k ★
- Optimization — ai-research-skills AWQ, GPTQ, GGUF, Flash Attention, bitsandbytes 5.4k ★
- CodeIsland — Real-time AI coding agent status panel in your MacBook notch — live status, approvals & replies for 13 AI tool 2.3k ★
README
mac code
**Run models that don't fit in RAM on your Mac. $0/month.**
Can I run this on my Mac?
| Your Mac | RAM | What you can run | Speed |
|---|---|---|---|
| Any Mac | 8 GB | Qwen3.5-9B (Q4_K_M, 5.3 GB), 4K context | 16-20 tok/s |
| Any Mac | 16 GB | Qwen3.5-9B (Q4_K_M, 5.3 GB), 64K context | 16-20 tok/s |
| Mac mini M4 | 16 GB | Qwen3.5-35B-A3B (IQ2_M, 10.6 GB) | 30 tok/s |
| Mac mini M4 | 16 GB | Qwen3-30B-A3B Q4 (17.2 GB) via Expert Sniper | 4.3 tok/s |
| Mac mini M4 | 16 GB | Qwen3.5-35B-A3B Q4 (19.5 GB) via Expert Sniper | 5.4 tok/s |
| Mac mini M4 | 16 GB | Qwen3.5-35B-A3B Q4_K_M (22 GB) via Flash Streaming | 1.54 tok/s |
| Mac mini M4 | 16 GB | Qwen3.5-27B (16.1 GB) via Flash Streaming | 0.18 tok/s |
| Mac mini M4 Pro | 48 GB | 35B at full Q4 in RAM | 30+ tok/s |
**"I wanted to run the Qwen 27B on my M2 16GB but failed. That's not possible, right?"**
It is possible. We stream FFN weights from SSD — only 5.5 GB stays in RAM. The output is coherent, full 4-bit quality. It's slow (0.18 tok/s on a Mac mini M4) but the method works on any 16 GB Apple Silicon Mac. No 2-bit compression, no mmap thrashing, no swap death. [See how it works.](#how-flash-streaming-works)
Quick Start
35B Agent (recommended — 30 tok/s on 16 GB)
The fastest option. Uses llama.cpp with a 2-bit quantization (IQ2_M) that fits entirely in RAM.
brew install llama.cpp
pip3 install rich ddgs --break-system-packages
# Download model (10.6 GB)
python3 -c "
from huggingface_hub import hf_hub_download
hf_hub_download('unsloth/Qwen3.5-35B-A3B-GGUF',
'Qwen3.5-35B-A3B-UD-IQ2_M.gguf', local_dir='$HOME/models/')
"
# Start server + agent
llama-server \
--model ~/models/Qwen3.5-35B-A3B-UD-IQ2_M.gguf \
--port 8000 --host 127.0.0.1 \
--flash-attn on --ctx-size 12288 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--n-gpu-layers 99 --reasoning off -np 1 -t 4
python3 agent.
Related Skills
Imagegen Mac
Local AI image generation for Mac. Run Qwen-Image 2.1 on Apple Silicon, offline and private, with a built-in M
LLM Launchpad
一条命令在 Apple Silicon Mac 上部署本地大模型(Qwen3.8-27B · Ollama · Claude Code 兼容)
Setup Lmstudio VSCode
Set up a fully local coding model in VS Code on an Apple Silicon Mac (LM Studio MLX engine, no admin), with au
Vllm Mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuou
Wax
Shared Single-file memory layer for all your agents, sub mili-second RAG over text, photo and video on Apple S
Ovo Local LLM
A private Claude-Code-style coding agent for Apple Silicon — run chat, code, and local model workflows on-devi
Related Agents
Buildhost Lead
Use PROACTIVELY and automatically — do not wait to be asked — as the Claude-Code node lead FOR an Apple-silico
Merchandiser
Owns BlushTip's product line and launch calendar: sets and bundles, the price ladder around the $50 free-shipp
Mcg
Product management author and coach. Founder and partner of Silicon Valley Product Group (SVPG, 2001). Author