Agent Eval Design
Description
--- description: "Design an evaluation for an AI agent or LLM feature: what to test, how to grade it, and how to catch regressions." argument-hint: "[system]" --- You were invoked as a slash command. The user's input: $ARGUMENTS Use that input to fill this prompt's variables (take the main content, topic, or task from it; ask only if a required value is missing and not supplied), then follow the prompt exactly. --- Design an evaluation for: {system} Known/feared failures: {failures} Build
Installation
Installs to ~/.claude/skills/amey-thakur-ai-skills-agent-eval-design/SKILL.md
mkdir -p ~/.claude/skills/amey-thakur-ai-skills-agent-eval-design && curl -fsSL https://raw.githubusercontent.com/Amey-Thakur/AI-SKILLS/HEAD/commands/agent-eval-design.md -o ~/.claude/skills/amey-thakur-ai-skills-agent-eval-design/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Skills
Spec Kit
💫 Toolkit to help you get started with Spec-Driven Development
Testing Webapp Testing
Test local web applications using Playwright for UI verification and debugging
Testing #29
, [#52](https://github.com/affaan-m/everything-claude-code/issues/52), [#103](https://github.com/affaan-m/ever
Testing Fix Issue
by metabase - Addresses GitHub issues by taking issue number as parameter, analyzing context, implementing sol
Testing Pypict Test Design
Design comprehensive test cases using PICT (Pairwise Independent Combinatorial Testing) for optimized test sui
Testing gstack
| 15,000+ | Garry Tan's exact Claude Code setup: 6 opinionated tools that serve as CEO, Eng Manager, Release M
Testing