AI Eval Engineer — Data & AI agent for Claude Code
Builds and maintains the eval pipeline for ai-system / agent-product archetypes.
How to install AI Eval Engineer
Installs to ~/.claude/agents/avelikiy-great_cto-ai-eval-engineer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/avelikiy/great_cto/HEAD/agents/ai-eval-engineer.md -o ~/.claude/agents/avelikiy-great_cto-ai-eval-engineer.md Restart Claude Code, or start a new session, for it to be picked up.
What AI Eval Engineer does
name: ai-eval-engineer description: Builds and maintains the eval pipeline for ai-system / agent-product archetypes. Outputs tests/eval/EVAL-*.md files (golden citation, refuse-when-uncertain, output schema, prompt injection, cost-overrun, cross-user isolation). Runs regression on every prompt or model change. Detects drift. model: haiku tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, memory_20250929, advisor_20260301 maxTurns: 30 timeout: 600 effort: MEDIUM memory: project
Alternatives in Data & AI
- LLM AI Hunter — LLM and Agentic AI vulnerability specialist 812 ★
- LLM Expert — Use this agent for AI SDK integration, LLM provider configuration, prompt template management, error handling 434 ★
- API Cost Scanner — Scans codebase for AI/API cost optimization opportunities including model usage, token estimates, caching stra 160 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Encounter Designer
Game encounter / combat designer sub-agent. Produces enemy archetypes, boss phases, encounter pacing, AI tunin
Askit Quality Grader
Judges whether a skill triggers and behaves correctly by running it against its eval-set and grading the outpu
Sqli Hunter
Active SQL injection hunter for an ingested program. Consumes webvuln-surface injection points + auth-context,
AI Ethics Reviewer
AI / ML ethics + responsible-AI specialist — bias, fairness, model selection, dataset provenance, automated-de
Prompt Polisher
Analyzes AI prompt files (.claude/agents/.md, .claude/skills//.md, .agents/skills//.md, CLAUDE.md, AGENTS.md,
Eval Inline Fact
Eval arm inline-fact (Q4 — the CALCULATED vs MEASURED line reduced to its FACT half: what the quantities are,
Related Skills
Claude Ecosystem Wiki
An agent-maintained Obsidian wiki about the Claude ecosystem. Claude Code writes and maintains every page unde
Frank
Activates Frank, DevOps Engineer for production infrastructure. Monitors health, manages deployments, handles
Eval Author
Draft a new eval case (eval/cases/ .json) + scaffold its golden folder, golden_pending until captured on the V