Evo — Development skill for Claude Code
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
How to install Evo
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open evo-hq/evo and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Evo does
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
Alternatives in Development
- Research Codebase — You are tasked with conducting comprehensive research across the codebase to answer user questions by spawning 10k ★
- Tentacle Planner — You are the Tentacle Planner — a meta-agent that analyzes this codebase and creates department tentacles to or 1.4k ★
- Agent Skill — A Claude Code plugin marketplace containing the ast-grep skill for powerful structural code search using Abstr 851 ★
README
evo
[](https://pypi.org/project/evo-hq-cli/) [](LICENSE) [](https://github.com/evo-hq/evo/actions/workflows/ci.yml) [](https://doi.org/10.5281/zenodo.20447923)
**Get started with autoresearch on any codebase - with two simple commands.**
Do you want to do more with autoresearch or need a custom, hands-on deployment? [Request access to evo platform](https://evo-hq.com/beta) or email [hello@evo-hq.com](mailto:hello@evo-hq.com).
**[Try it](#try-it)** · **[Install](#install)** · **[How it works](#how-it-works)** · **[Dashboard](#dashboard)** · **[Upgrading](#upgrading)**
You give it a codebase. It discovers metrics to optimize, sets up the evaluation, and starts running experiments in a loop -- trying things, keeping what improves the score, throwing away what doesn't.
*Inspired by [Karpathy's autoresearch](https://github.com/karpathy/autoresearch)* -- where an LLM runs training experiments autonomously to beat its own best score. Autoresearch is a pure hill climb: try something, keep or revert, repeat on a single branch. Evo adds structure on top of that idea:
- Tree search over greedy hill climb. Multiple directions can fork from any committed node, so exploration doesn't collapse to one path.
- Parallel semi-autonomous agents. Spawn multiple subagents and run them simultaneously, each in its own git worktree. Each subagent reads traces, formulates hypotheses, and can run multiple iterations within its branch.
- Shared state. Failure traces, annotations, and disc
Related Skills
Ranking Eval Loop
Ranking eval loop — mine real misses, then measure the fix
Adw Eval
Measure the loop itself, then PROPOSE (never silently apply) up to 3 improvements
Hatch3r Onboard
Generate a comprehensive onboarding guide for a new developer joining the project -- spawn parallel researcher
Forge Benchmark
Measure validation posture across 5 dimensions with trend tracking
Benchmark Rerankers
Measure whether adding a reranker actually improves retrieval, by scoring reranked vs. un-reranked results on
Vault Autoresearch
3-round autonomous research loop with intermediate notes and synthesis
Related Agents
Autoresearch Orchestrator
Use this agent to run an autoresearch experiment session end-to-end — setup, baseline, and a batch of edit→mea
Release Agent
Use when releasing a new version. Discovers how the repository releases, then performs the release.
Deal Sourcer
Builds and qualifies the acquisition target list. Maps a sector, discovers companies matching the investment t