Eval Grader — Development agent for Claude Code
Use when an eval case's subject run has finished and needs an independent verdict against its expectations.
How to install Eval Grader
Installs to ~/.claude/agents/hodamousavipour-claude-code-llm-engineering-eval-grader.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/hodamousavipour/claude-code-llm-engineering/HEAD/agents/eval-grader.md -o ~/.claude/agents/hodamousavipour-claude-code-llm-engineering-eval-grader.md Restart Claude Code, or start a new session, for it to be picked up.
What Eval Grader does
name: eval-grader description: Use when an eval case's subject run has finished and needs an independent verdict against its expectations. Spawned by scripts/run_eval.py via `claude -p --agent eval-grader` after each subject run, non-interactively, on a model pinned distinct from the subject's — never invoke this agent to modify the code, skill, or transcript under review, only to judge it. model: opus tools:
- Read
- Grep
- Glob
You are an independent judge grading a completed ev
Alternatives in Development
- Sales Automator — Draft cold emails, follow-ups, and proposal templates 2.8k ★
- Engram Assessor — Independent grader of learner productions for the Engram learning plugin 1.4k ★
- Eval Implementer — Implementer role of the D365FO agent eval loop 136 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Harness Prover
Running-app verifier for harness eval loops. Drives the live feature (browser, API, or CLI) and returns a bina
Aidd Independent Thinker
AIDD devil's advocate: the strongest honest case against the chosen architecture — counter-arguments, hidden a
Stroi Validator
Use this agent when a change needs an independent pass/fail verdict rather than an opinion. Typical triggers i
Skill Eval Grader
스킬 eval 용 채점자 — 러너 결과를 assertion 리스트로 채점하고 엄격한 JSON 판정을 반환한다. 채점과 동시에 assertion 품질(변별력) 자체도 비평한다. skill-creato
Proof Auditor
Independent rubric-verdict producer for reasoning deliverables. Runs alongside the incumbent judge; produces a
Candidate
Answers interview questions in first person AS Max, grounded strictly in his real experience. Spawned by /mock
Related Skills
Finalise
Familiar: Title, framing, subject line and SEO, once the piece is finished
Tldr Commit
Generate tldr-style commit message (verdict first, ≤50 char subject, why over what)
Cmd Judge
Adversarial verifier for the PROVE phase: treats a finished-work report as untrusted claims, re-runs every ver