Skill Eval Reporter — Development agent for Claude Code
Compares repeated paired execution results using blind A/B methodology and generates a skill effectiveness report.
How to install Skill Eval Reporter
Installs to ~/.claude/agents/shinpr-rashomon-skill-eval-reporter.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/shinpr/rashomon/HEAD/agents/skill-eval-reporter.md -o ~/.claude/agents/shinpr-rashomon-skill-eval-reporter.md Restart Claude Code, or start a new session, for it to be picked up.
What Skill Eval Reporter does
name: skill-eval-reporter description: Compares repeated paired execution results using blind A/B methodology and generates a skill effectiveness report. Use when valid skill-evaluation result pairs are available. tools: Read skills: prompt-optimization
You are a specialized agent for evaluating skill effectiveness through blind comparison.
Initial Mandatory Task
Read `prompt-optimization/references/execution-quality.yaml` and `prompt-optimization/references/skills.md`. Use the fir
Alternatives in Development
- Eval Engineer — GAIA evaluation framework specialist 1.5k ★
- Skill Effectiveness Analyzer 271 ★
- Coder Critic — Code critic that reviews R/Python/Julia scripts for strategic alignment, code quality, numerical discipline, a 250 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Learner Diagnostic
Assesses learning approach across 5 dimensions from Justin Sung's methodology. Use when user wants to evaluate
Player Eval
Screenshot evaluation agent. Analyzes one screenshot and outputs structured JSON perception/decision.
Disc Reviewer Ambiguity
The ambiguity judge of the stage-1 discovery review — compares two blind readers' builds and reports every rea
Blind Evaluator
Structurally separate eval agent. Receives ONLY the problem statement + rubric, NEVER the solution or the impl
Blind Probe Author
Writes efficacy-eval scenarios, retrieval probes, and held-out criteria for a document type WITHOUT reading th
Git Hygiene Reviewer
Reviews the shape of the history and the PR meta, never the code. Atomic conventional commits, linear bisectab
Related Skills
Adhd Output Style
AUDHD output style for Claude Code + the blind paired eval harness that gates style changes on correctness
Agent PR Replay
Agent PR Replay takes merged PRs from any repository, reverse-engineers the task prompt, runs Claude Code agai
Meta Review
Run a quarterly framework evaluation (macro loop). Assesses agent effectiveness, architectural drift, rule upd