Evaluate — Testing & QA agent for Claude Code
Evidence-based auditor in a fresh context.
How to install Evaluate
Installs to ~/.claude/agents/juanwmedia-craft-evaluate.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/juanwmedia/craft/HEAD/agents/evaluate.md -o ~/.claude/agents/juanwmedia-craft-evaluate.md Restart Claude Code, or start a new session, for it to be picked up.
What Evaluate does
name: evaluate description: Evidence-based auditor in a fresh context. Verifies every claim in one named artifact (a spec cut, a board's coverage claims, a graduation list, a plan, a report) against the primary sources and returns a verdict per claim with file:line evidence. Never edits anything. model: sonnet tools: Read, Grep, Glob, Bash, WebFetch, WebSearch
You are the evaluator half of the Evaluator-Optimizer pattern: a ruthless, evidence-based auditor of the artifact the caller nam
Alternatives in Testing & QA
- Test Wiring Auditor — 変更差分に対してテスト網が追随しているかを fresh-context で監査する read-only auditor 3.1k ★
- Fable Verifier — Cold verifier for finished fable deliverables 837 ★
- Witness — Independently witness that an Allium loop's convergence claim is true and was reached honestly 474 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
MCP Currency
MCP protocol & Kotlin SDK currency expert. Researches the latest MCP spec revision, kotlin-sdk-server releases
Vnv Sprint Prep Agent.Agent
Runs the V&V sprint workflow — platform-parity check, cloning dev sprint stories to the V&V board, validating
Feature Spec
Turns a rough feature idea into a tight, buildable specification with scope, acceptance criteria, and an expli
QA Correctness
Broken-output QA gate. Verifies text is present and not cut off, no black frames, no overlap, no stuck/duplica
Forge Visual Verifier
Perceptual gate for spec [visual] acceptance criteria. Drives Playwright MCP (navigate + take_screenshot + eva
Evolve Behavior Baseline
Behavior baseline agent for the Evolve Loop (Evaluate archetype). The advisor INSERTS this phase on refactor c
Related Skills
Deep Ass Research
IDE-agnostic deep-research framework for AI agents: maps a topic wide, commits to depth on purpose, adversaria
Reviewer Gap Audit
You are the compact global omission auditor for a decomposed pull-request review. Primary specialists already
Sharpen
Cut hedges and weasel words; commit to claims without inflating them