NiyaNagi

Claude Token Stack — AI skill for Claude Code

AI community

Accuracy-guarded token-efficiency stack for Claude Code on Windows & the Claude desktop app: code graph, context sandbox, RTK, guarded Headroom, caveman, subagent model router, compaction handoff, ski.

How to install Claude Token Stack

This entry records only its repository, not the path inside it, so there is no exact command to give. Open NiyaNagi/claude-token-stack and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Claude Token Stack does

Accuracy-guarded token-efficiency stack for Claude Code on Windows & the Claude desktop app: code graph, context sandbox, RTK, guarded Headroom, caveman, subagent model router, compaction handoff, skill intake.

Alternatives in AI

  • Fusion Plan — OMC interview (interactive) → 3-round Opus 4.8 + GPT-5.5 fusion deepening → concise .omc/plans/ output → revie 468 ★
  • Snip — CLI proxy that reduces LLM token usage by 60-90% 425 ★
  • Project Butler — Project memory system for AI coding assistants (Claude Code, Cursor, Codex): session logs, project wiki, rules 372 ★

README

claude-token-stack

A layered, **accuracy-guarded** token-efficiency setup for Claude Code on Windows — including Claude Code running inside the **Claude desktop app** (Code tab), where proxy-based tools don't work.

It combines five open-source tools with one custom hook dispatcher (`ts.js`) that enforces them, guards them against silent information loss, routes subagents to the right model, survives compaction with a handoff, and puts every new skill through an intake audit.

Inspired by *"How I Cut Claude Code Token Usage by 90%+ With 5 Tools, Custom Hooks, and Enforcement"* (sgaabdu4/claude-code-tips, Apr 2026). This repo is a Windows/desktop-app port that was tested live, with several of that article's settings corrected and extra safeguards added where testing showed real accuracy loss. See [What's different from the article](#whats-different-from-the-article).

**Quick start:** [`BOOTSTRAP_PROMPT.md`](BOOTSTRAP_PROMPT.md) — paste one prompt into Claude Code and it installs, verifies and benchmarks everything with you. Or run `install.ps1` yourself ([docs/SETUP.md](docs/SETUP.md)).


Architecture

flowchart TD
  P[Prompt] --> UPS[UserPromptSubmit: caveman mode, 'exact' toggle]
  UPS --> M[Main model]
  M -->|Agent call| R[agent-route: pick haiku/sonnet/opus, learn from escalations]
  M -->|Read / Grep source| G[cbm-gate: query code graph first]
  G --> CBM[(codebase-memory-mcp graph)]
  M -->|Bash / PowerShell| S[shell-gate: ban raw dumps, RTK rewrite except diffs]
  S --> OUT[tool output]
  OUT --> C[compress: Headroom on huge repetitive shell output, signal lines verified]
  M -->|big analysis| CTX[context-mode sandbox: only summaries return]
  M --> CAVE[caveman lite: terser replies]
  PC[PreCompact] --> H[handoff file] --> SS[SessionStart: re-inject handoff, list pending skill intakes]
Layer Tool What it saves Where
1 codebase-memory-mcp (cbm)