LLM Diet — AI skill for Claude Code
Cut AI coding tokens by 99%.
How to install LLM Diet
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open ShresthSamyak/LLM_DIET and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What LLM Diet does
Cut AI coding tokens by 99%. Deterministic context injection for Claude Code, Cursor, Windsurf.
Alternatives in AI
- Jcodemunch MCP — Cut AI token costs 95%+ on code exploration 2.6k ★
- Codesight — Universal AI context generator 979 ★
- Skilldock — SkillDock is an AI skill manager and skill management desktop app for Claude Code, Cursor, Codex, Windsurf, Ge 490 ★
README
llm-diet
**Give Claude the right context upfront. Fewer turns, faster answers, lower cost.**
[](https://pypi.org/project/llm-diet/) [](LICENSE) [](https://pypi.org/project/llm-diet/)
Deterministic context retrieval for AI coding tools. Parses your repo into a call graph, intercepts every file read Claude makes, and returns compressed versions — so Claude explores freely but cheaply.
The Problem
Every Claude Code session starts blind. Claude explores your entire codebase before answering — reading files, listing directories, running commands. That exploration costs tokens and time.
Without llm-diet:
Claude reads 10 files × 8,000 tokens = 80,000 tokens consumed
Cost: $0.19 for a simple bug fix session
With llm-diet:
Claude reads 10 files × 300 tokens = 3,000 tokens consumed
Cost: $0.025 for the same session
How It Works
User prompt
│
▼
context-engine (call graph)
│ scores every function against your query
▼
Claude Code session opens
│
▼
Claude calls read_file("validators/amazon.py")
│
▼
llm-diet-shadow MCP server intercepts
│ returns compressed 872-token version
│ instead of raw 6,590-token file
▼
Claude answers — correctly — using compressed context
Claude thinks it explored. It did — but every read returned our compressed version, not the raw file.
Benchmark
**Tested on coupon-hunter-poc (40-node Python project)**
| File | Original | Compressed | Reduction |
|---|---|---|---|
| validators/playwright_amazon.py | 6,590 chars | 872 chars | 86% |
| orchestrator.py | 10,492 chars | 2,169 chars | 79% |
| connectors/playwright_amazon.py | 3,067 chars | 631 chars | 79% |
| openrouter_agent.py | 2,860 chars | 966 chars | 66% |
| retailmenot_scraper.py | 2,705 chars | 960 chars |
Related Skills
Toknt
Tokens? Tokn't. Local-first token optimization for AI coding agents (Cursor, Claude Code, Codex). Cut redundan
Agent Context Framework
Production-ready Markdown (.md) context templates for AI coding agents (Cursor, Claude Code, Aider, Windsurf).
Context Broker
Cut Claude Code token usage 80-95% on big-file reads. Deterministic, ~1ms, zero egress — no second LLM.
Tokenomy
Tokenomy — A surgical toolkit for AI coding CLIs. Cut tokens, keep context, build faster.
Rustygrep
Fast grep with token-compressed, AI-native output for LLM coding agents. --llm cuts context-window tokens 60-9
Smart Context Delegation
Cut coding-agent token usage by ~90%: a PreToolUse hook blocks large file reads; bulk-read/code-write delegate
Related Agents
10x Tool Calls
Cursor and Windsurf meter usage by requests and tool calls rather than tokens, which means a finished or stall
Context Auditor
Use proactively when the user wants to cut Claude Code token cost or asks "why is this session so expensive /
LLM Sec Review
LLM/agent security review specialist — prompt injection, the Agents Rule of Two, tool-call authorization, mode