Ai Eval Designer
Description
--- name: ai-eval-designer description: Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design evals for", "I need an eval set", "what should I test for this AI feature", "build a golden dataset", or before launching an AI feature that doesn't have evals yet. Also trigger after an incident, to add the new failure mode as eval cases. tools: Read, Grep, Glob, Write model: opus last-reviewed: 2026-05-17 skills: - ai-playbook --- # AI Eval Desi
Installation
Installs to ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/kumaran-is/claude-code-onboarding/HEAD/.claude/agents/ai-eval-designer.md -o ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Agents
Gitnexus Test Ci Verifier
GitNexus test and CI reviewer. Use to verify whether changed behavior is covered by targeted tests, whether CI
Testing & QA Track Plan
| **Track ID:** {{TRACK_ID}} **Spec:** [spec.md](./spec.md) **Estimated Effort:** {{EFFORT_ESTIMATE}}... | - |
Testing & QA Track Spec
| **Track ID:** {{TRACK_ID}} **Type:** {{TRACK_TYPE}} (feature | bug | chore | refactor) **Priority:**... | -
Testing & QA Workflow
| 1. **plan.md is the source of truth** - All task status and progress tracked in the plan 2. **Test-D... | -
Testing & QA Test Generate
| You are a test automation expert specializing in generating comprehensive, maintainable unit tests a... | -
Testing & QA Claude Code Flow
by ruvnet - This mode serves as a code-first orchestration layer, enabling Claude to write, edit, test, and op
Testing & QA