Eval Engineer — Development agent for Claude Code
GAIA evaluation framework specialist.
How to install Eval Engineer
Installs to ~/.claude/agents/amd-gaia-eval-engineer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/amd/gaia/HEAD/.claude/agents/eval-engineer.md -o ~/.claude/agents/amd-gaia-eval-engineer.md Restart Claude Code, or start a new session, for it to be picked up.
What Eval Engineer does
name: eval-engineer description: GAIA evaluation framework specialist. Use PROACTIVELY for writing eval tests, generating ground truth, running batch experiments, benchmarking models, or analysing transcripts. tools: Read, Write, Edit, Bash, Grep, Glob model: opus
You work on GAIA's evaluation framework in `src/gaia/eval/`. Your job is making model/agent comparisons reproducible.
Output style
Follow [`CLAUDE.md`](../../CLAUDE.md) → "How You Communicate".
NEVER run evals in para
Alternatives in Development
- Leader — Central decision-maker that plans experiments and reflects on results 728 ★
- Bench — API performance benchmarking — latency profiling, throughput testing, performance regression detection 73 ★
- Prototyper — The Prototyper builds rapid proof-of-concept implementations, throwaway spikes, and time-boxed experiments to 71 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Sme Eval Triage
Triage a golden-set failure from the compliance-SME seat before anyone edits ground truth. Use whenever make e
Scientific Figures
Generate publication-quality scientific figures with consistent styling. Use when experiments need visualizati
Sonmat Witness
External witness agent. Verifies intent-artifact match using user turn cascade and ground truth. Protocol-isol
Stroi Explorer
Use this agent when a planner or skill needs ground truth about the codebase before designing against it. Typi
Doc Analyser
Per-page documentation analyser for the ground-truth doc-lineage layer. Reads one published doc page (the GitB
Cutter
Turn recorded biographies into the shaped scene-list of a novel — the documentary editor's cut. Given the exha
Related Skills
Ground Truth Evals
An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models
Rails AI Context
45 MCP tools that give AI coding agents ground truth about your Rails app: schema, models, routes, controllers
Eval Audit
/eval-audit — Batch Skill Evaluation