Eval Failure Diagnoser
Description
--- name: eval-failure-diagnoser description: Diagnoses promptfoo skill-eval failures for sumo-qa. Use after `npm run eval` or `npm run eval:all` returns any FAIL. Reads the eval output, identifies which assertion failed (shape / grounding / anti-pattern / javascript), locates the relevant SKILL.md section, and recommends strengthening the skill — never loosening the rubric. Returns a structured diagnosis per failure. Does NOT edit SKILL.md or YAML files. tools: Bash, Read, Grep, Glob --- # eva
Installation
Installs to ~/.claude/agents/sumithr-sumo-qa-eval-failure-diagnoser.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/sumithr/sumo-qa/HEAD/.claude/agents/eval-failure-diagnoser.md -o ~/.claude/agents/sumithr-sumo-qa-eval-failure-diagnoser.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Agents
Gitnexus Test Ci Verifier
GitNexus test and CI reviewer. Use to verify whether changed behavior is covered by targeted tests, whether CI
Testing & QA Track Plan
| **Track ID:** {{TRACK_ID}} **Spec:** [spec.md](./spec.md) **Estimated Effort:** {{EFFORT_ESTIMATE}}... | - |
Testing & QA Track Spec
| **Track ID:** {{TRACK_ID}} **Type:** {{TRACK_TYPE}} (feature | bug | chore | refactor) **Priority:**... | -
Testing & QA Workflow
| 1. **plan.md is the source of truth** - All task status and progress tracked in the plan 2. **Test-D... | -
Testing & QA Test Generate
| You are a test automation expert specializing in generating comprehensive, maintainable unit tests a... | -
Testing & QA Claude Code Flow
by ruvnet - This mode serves as a code-first orchestration layer, enabling Claude to write, edit, test, and op
Testing & QA