Evolve Benchmark Gate — Development agent for Claude Code
Statistical benchmark comparison gate (Evaluate archetype).
How to install Evolve Benchmark Gate
Installs to ~/.claude/agents/mickeyyaya-evolve-loop-evolve-benchmark-gate.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/mickeyyaya/evolve-loop/HEAD/agents/evolve-benchmark-gate.md -o ~/.claude/agents/mickeyyaya-evolve-loop-evolve-benchmark-gate.md Restart Claude Code, or start a new session, for it to be picked up.
What Evolve Benchmark Gate does
name: evolve-benchmark-gate description: Statistical benchmark comparison gate (Evaluate archetype). model: tier-2 capabilities: [file-read, search, shell, file-write] tools: ["Read", "Grep", "Glob", "Bash", "Write"] perspective: "statistical-performance-gatekeeper" output-format: "benchmark-gate-report.md"
Evolve Benchmark Gate Agent
You are the **Benchmark Gate** agent in the Evolve Loop. Your job is to run a statistical benchmark comparison of the modified code against a stored ba
Alternatives in Development
- Livemsg Gate — セッション間メッセージの主張を裏取りして SEND / HOLD を返す read-only gate 3.1k ★
- Release Gate Runner 2.1k ★
- Dotnet Benchmark Designer 874 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Evolve Behavior Compare
Behavior comparison agent for the Evolve Loop (Evaluate archetype). The advisor INSERTS this phase on refactor
Evolve Adversarial Review
Adversarial-review agent for the Evolve Loop (Evaluate archetype). The advisor INSERTS this phase after build
Crucible Qg Judge
Stagnation judge for Crucible quality-gate — mechanical cross-round comparison of finding sets. Pinned to Sonn
Decision Gate
Takes the comparison table and classifies the final 1–3 ETF candidates as worth reviewing / conditional / on h
Verify Logic
A verification agent that performs Stage 2 verification — comparing the tables and figures embedded/referenced
Evidence Reviewer
Skeptical auditor of evidence bundles against .claude/skills/evidence-standards.md. Detects circular citations
Related Skills
Refresh Coder Comparison
Re-research the coder capability matrix — re-check benchmark boards, pricing, model rosters, safety evaluation
Bench Run
Run a benchmark comparison on the named scope. Without a scope, list available ones and ask.
Cceval
Evaluate and benchmark your CLAUDE.md effectiveness with automated testing