kumaran-is

AI Eval Designer — Testing & QA agent for Claude Code

Testing & QA community

Use this agent to design a risk-tiered evaluation set for an AI feature.

How to install AI Eval Designer

Installs to ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md

Terminal
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/kumaran-is/claude-code-onboarding/HEAD/.claude/agents/ai-eval-designer.md -o ~/.claude/agents/kumaran-is-claude-code-onboarding-ai-eval-designer.md

Restart Claude Code, or start a new session, for it to be picked up.

What AI Eval Designer does


name: ai-eval-designer description: Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design evals for", "I need an eval set", "what should I test for this AI feature", "build a golden dataset", or before launching an AI feature that doesn't have evals yet. Also trigger after an incident, to add the new failure mode as eval cases. tools: Read, Grep, Glob, Write model: opus last-reviewed: 2026-05-17 skills:

  • ai-playbook

AI Eval Desi

Alternatives in Testing & QA

  • Pencil UI Designer — Use this agent when you need Pencil-based design workflows: atomic designer actions, MCP tools, and spec-drive 635 ★
  • Module Test Writer — Use this agent to write feature tests for LaraDashboard module CRUD operations, services, policies, and Livewi 407 ★
  • Playwright Test Generator — Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants 246 ★

Full documentation available on GitHub

View Source Repository