Eval Coach — AI skill for Claude Code
Agent Skill for Evaluation-Driven Development (EDD) - guide AI evaluation strategies with the 50-40-10 rule.
How to install Eval Coach
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open BayramAnnakov/eval-coach and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Eval Coach does
Agent Skill for Evaluation-Driven Development (EDD) - guide AI evaluation strategies with the 50-40-10 rule.
Alternatives in AI
- Bmad Method — Breakthrough Method for Agile Ai Driven Development 42712 382 2 42.9k ★
- Claude Task Master — by eyaltoledano - A task management system for AI-driven development with Claude, designed to work seamlessly 26k ★
- Hindsight — State-of-the-art long-term memory for AI agents by Vectorize 6.7k ★
README
Eval Coach
An Agent Skill for designing comprehensive AI evaluation strategies using Evaluation-Driven Development (EDD).
About
Created by [Bayram Annakov](https://linkedin.com/in/bayramannakov) while building [Onsa.ai](https://onsa.ai) - AI agents for B2B sales prospecting.
If you find this useful, say hi on LinkedIn!
What is Eval Coach?
Eval Coach guides you through a structured 5-step framework for evaluating LLM applications:
- Define Success - Map business goals to measurable metrics
- Design Dataset - Create diverse test cases (happy path, edge cases, adversarial)
- Select Methods - Choose Automated, LLM-as-Judge, or Human evaluation
- Plan Automation - Integrate evals into CI/CD
- Monitor Production - Track drift and collect feedback
Installation
Claude Code / Cursor / VS Code
Copy this skill to your project's skills directory:
git clone https://github.com/BayramAnnakov/eval-coach.git ~/.claude/skills/eval-coach
Or add to your project:
mkdir -p skills
git clone https://github.com/BayramAnnakov/eval-coach.git skills/eval-coach
Usage
Invoke the skill by name or with trigger keywords:
/eval-coach
Or just mention evaluation-related topics:
- "help me create an evaluation strategy"
- "design test cases for my agent"
- "set up LLM testing"
The 50-40-10 Rule
A practical distribution for evaluation methods:
| Tier | Method | Cost | Percentage |
|---|---|---|---|
| 1 | Automated (schema, keywords, latency) | $0.00/run | 50% |
| 2 | LLM-as-Judge (quality, relevance) | $0.01-0.05/run | 40% |
| 3 | Human Review | $5-50/run | 10% |
Templates Included
templates/dataset.py- LangSmith dataset creation with example test casestemplates/evaluators.py- 10 ready-to-use evaluators (automated, LLM-as-Judge, performance)templates/compare.py- Experiment comparison utilities
Key Insight: Silent Failures
The most dangerous failure
Related Skills
Agent Belt
Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor
Buyer Eval Skill
B2B software vendor evaluation skill for Claude Code — domain-expert questions, vendor AI agent conversations,
Agent Eval Workbench
MLflow-based evaluation kit for AI agents: golden-task runner, calibrated LLM judge, list-price cost accountin
Bullpen Coach
Switch Sage (the wellness coach) on or off. When on, Sage sends proactive check-ins (rest reminders, ship cele
Study Coacher
AI study coach skill: spaced repetition, retrieval practice & deliberate drills as a daily AI-driven loop — wi
Learn Skill
Agent skill that distills AI study chats into a causal-chain learning wiki — mechanical checks, a regression e
Related Agents
DataAnalyst
Data analyst — success metrics, KPIs, measurement strategies, business impact analysis, data-driven evaluation
TDD Coach
Test-driven development guidance, test writing, coverage improvement
Pi Coach En
PI Wisdom-in-Action Coach v20 — monitors teammate progress with classical Chinese wisdom and MBTI cognitive st