EvalPlan — AI skill for Claude Code
Phase 3 (+ai) — build golden / adversarial / regression eval sets, graders, thresholds, baseline.
How to install EvalPlan
Installs to ~/.claude/skills/linofcp007-dev-spec-driven-evalplan/SKILL.md
mkdir -p ~/.claude/skills/linofcp007-dev-spec-driven-evalplan && curl -fsSL https://raw.githubusercontent.com/linofcp007/dev-spec-driven/HEAD/commands/evalPlan.md -o ~/.claude/skills/linofcp007-dev-spec-driven-evalplan/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What EvalPlan does
description: Phase 3 (+ai) — build golden / adversarial / regression eval sets, graders, thresholds, baseline. PT - plano de evals (+ai). ES - plan de evals (+ai). argument-hint: "[feature name]"
Use the **dev-spec-driven** skill, Phase 3 (Eval Plan). Only relevant when the feature is on the **+ai** track.
Feature: $ARGUMENTS
Write `eval-plan.md`: a **golden** set (50–200 representative inputs with expected quality), an **adversarial** set (prompt injections, jailbreaks, out-of-scope,
Alternatives in AI
- Codex — openai-codex Adversarial code review, Codex CLI integration, cross-model analysis 9.6k ★
- Internal Safety Collapse — We built an adversarial codespace setup 1.2k ★
- UltraCode Shim — Give Claude Code's ultracode mode to ANY model you already pay for 422 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Claude Evals
Production eval framework for Claude Agent SDK — implements Anthropic's published eval patterns with native SD
Agent Eval Workbench
MLflow-based evaluation kit for AI agents: golden-task runner, calibrated LLM judge, list-price cost accountin
Harness Forge
One /forge command installs a self-improving Claude Code harness in any project: self-healing failure ledger,
Learn Skill
Agent skill that distills AI study chats into a causal-chain learning wiki — mechanical checks, a regression e
Ground Truth Evals
An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models
MigrateModel
(+ai) Eval-gated migration to a new/replacement model — never migrate blind. PT - migração de modelo com evals
Related Agents
AI Eval Designer
Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design ev
Sme Eval Triage
Triage a golden-set failure from the compliance-SME seat before anyone edits ground truth. Use whenever make e
Eval Author
Authoring role for the D365FO agent eval loop catalog. Drafts a new eval/cases/ .json spec (valid against eval