VLS 32M Context Local LLM — AI skill for Claude Code
Give Claude Code, Codex and opencode a 32M-token memory that never leaves your machine.
How to install VLS 32M Context Local LLM
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open Veloresearch/VLS-32M-Context-Local-LLM and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What VLS 32M Context Local LLM does
Give Claude Code, Codex and opencode a 32M-token memory that never leaves your machine. Local LLM server (MCP + OpenAI API) that runs any GGUF model with a context far larger than its window - no cloud, no retraining, no vector database. Windows and Linux, one command.
Alternatives in AI
- Optimization — ai-research-skills AWQ, GPTQ, GGUF, Flash Attention, bitsandbytes 5.4k ★
- Token Monitor — Local-first desktop widget for tracking token usage, costs, and limits across 32+ AI coding tools—including Cl 1.8k ★
- Paritok 4b V1 — Non-destructive compression gateway for AI coding agents 1.4k ★
README
[](https://github.com/Veloresearch/VLS-32M-Context-Local-LLM/releases/latest) [](https://github.com/Veloresearch/VLS-32M-Context-Local-LLM/releases) [](https://github.com/Veloresearch/VLS-32M-Context-Local-LLM) [](https://github.com/Veloresearch/VLS-32M-Context-Local-LLM)
[](#windows) [](#linux) [](LICENSE)
**[Install](#install)** · **[Velocity Context](#velocity-context)** · **[Everything in it](#everything-in-it)** · **[Mesh](#mesh)** · **[Benchmarks](#measured-not-claimed)** · **[API](#connect-anything)** · **[FAQ](#questions-people-actually-ask)**
**[veloresearch.com/vls](https://veloresearch.com/vls)**
A 4B model with an 8,19
Related Skills
Token Proxy
Local AI API gateway for OpenAI / Gemini / Anthropic. Runs on your machine, keeps tokens counted (SQLite), off
Token Counter
Counts tokens in documentation to ensure compliance with AI context window limits. Reports token counts per pa
Cordon
PII-redacting LLM compliance gateway — own your prompts; PII never leaves your perimeter
Axon
Local-first RAG engine with cited answers from your own docs. GraphRAG, sealed sharing, MCP-native, VS Code Co
TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Cl
Rails AI Context
45 MCP tools that give AI coding agents ground truth about your Rails app: schema, models, routes, controllers
Related Agents
Codex CLI Research Agent
Senior OpenAI researcher specializing in Codex CLI - researches latest context window and model availability
Openrouter Agent
Runs a model from any of ~60 vendors through OpenRouter's OpenAI-compatible endpoint, billed per token against
Cost Optimizer
Cloud and LLM cost optimization specialist — FinOps, right-sizing, caching strategies, Claude/OpenAI token red