AI Eval Designer — Testing & QA agent for Claude Code
Use this agent to design a risk-tiered evaluation set for an AI feature.
How to install AI Eval Designer
Installs to ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/kumaran-is/claude-code-onboarding/HEAD/.claude/agents/ai-eval-designer.md -o ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md Restart Claude Code, or start a new session, for it to be picked up.
What AI Eval Designer does
name: ai-eval-designer description: Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design evals for", "I need an eval set", "what should I test for this AI feature", "build a golden dataset", or before launching an AI feature that doesn't have evals yet. Also trigger after an incident, to add the new failure mode as eval cases. tools: Read, Grep, Glob, Write model: opus last-reviewed: 2026-05-17 skills:
- ai-playbook
AI Eval Desi
Alternatives in Testing & QA
- Pencil UI Designer — Use this agent when you need Pencil-based design workflows: atomic designer actions, MCP tools, and spec-drive 635 ★
- Module Test Writer — Use this agent to write feature tests for LaraDashboard module CRUD operations, services, policies, and Livewi 407 ★
- Playwright Test Generator — Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants 246 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Wh Designer
Design-phase agent. In one session turns " -> " into the risk tier, spec.md with ADRs, evals.md within the tie
Test Case Designer
Invoked after test-coverage-analyst has produced a gap report, or when the user says "design test cases for "
Sous Chef
Use this agent when reviewing Java/Spring Boot code, checking best practices, validating test coverage, or ens
AI Agent Eval Writer
Use this agent when you need to write evaluation tests for AI agents using Evalite or Autoevals frameworks. Th
QA Designer
Design test strategy and test cases from requirements — risk-based prioritization, coverage matrix, not test c
Eval Author
Authoring role for the D365FO agent eval loop catalog. Drafts a new eval/cases/ .json spec (valid against eval
Related Skills
Agent Eval Design
Design an evaluation for an AI agent or LLM feature: what to test, how to grade it, and how to catch regressio
Claude Evals
Production eval framework for Claude Agent SDK — implements Anthropic's published eval patterns with native SD
EvalPlan
Phase 3 (+ai) — build golden / adversarial / regression eval sets, graders, thresholds, baseline. PT - plano d