Hatch3r Benchmark — Development skill for Claude Code
Run and analyze performance benchmarks.
How to install Hatch3r Benchmark
Installs to ~/.claude/skills/hatch3r-hatch3r-hatch3r-benchmark/SKILL.md
mkdir -p ~/.claude/skills/hatch3r-hatch3r-hatch3r-benchmark && curl -fsSL https://raw.githubusercontent.com/hatch3r/hatch3r/HEAD/commands/hatch3r-benchmark.md -o ~/.claude/skills/hatch3r-hatch3r-hatch3r-benchmark/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Hatch3r Benchmark does
id: hatch3r-benchmark type: command orchestrator: true agentPipeline: [hatch3r-researcher, hatch3r-performance, hatch3r-docs-writer] description: Run and analyze performance benchmarks. Compare results against baselines, identify regressions, and produce performance reports. tags: [review, performance] quality_charter: agents/shared/quality-charter.md efficiency_patterns: agents/shared/efficiency-patterns.md cache_friendly: true parallel_tool_default: true efficiency_tier: standard triage_ti
Alternatives in Development
- Run Benchmarks — Run Benchmarks — GenAIIDP empirical config/scaling suite 295 ★
- Contract Review Skill — Review legal contracts for risks, extract key terms, and suggest redlines 192 ★
- Bib Coverage — Compare a project .bib against a Paperpile project/topic folder to find 130 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Warden Bench
Run the golden benchmark suite for one token-warden agent (or all) and compare results against the frozen run1
Gql Performance Benchmark
Run GraphQL performance benchmarks against the Alkemio API and compare against a stored baseline
Benchmark
Compare skill scores against ideal benchmarks
Forge Stats
Track command sizes, detect bloat, compare against baselines and thresholds
Sdlc Optimize
Performance optimization and profiling - identify bottlenecks, benchmark improvements
Hatch3r Onboard
Generate a comprehensive onboarding guide for a new developer joining the project -- spawn parallel researcher
Related Agents
Bench Runner
Executes a11y skill benchmarks across hosted and local model families. Runs cloud/Codex/Ollama benchmark scrip
Eval Failure Analyzer
Analyze Logic-Lens benchmark/eval failures. Use after running content-evals, or when pointed at a skills-works
SEO Drift
SEO drift analysis agent. Captures baselines of SEO-critical page elements and compares against stored snapsho