LLM Evaluator — Data & AI agent for Claude Code
LLM evaluation specialist.
How to install LLM Evaluator
Installs to ~/.claude/agents/g-hensley-gavins-agent-system-llm-evaluator.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/G-Hensley/gavins-agent-system/HEAD/agents/llm-evaluator.md -o ~/.claude/agents/g-hensley-gavins-agent-system-llm-evaluator.md Restart Claude Code, or start a new session, for it to be picked up.
What LLM Evaluator does
name: llm-evaluator description: LLM evaluation specialist. Use when designing eval suites for LLM features, calibrating LLM-as-judge prompts, interpreting drift, deciding pass thresholds. Dispatches alongside `ai-engineer` during build, and during incident review. Builds offline + online eval, tracing strategy, regression triggers on prompt changes. tools: Read, Write, Edit, Bash, Grep model: sonnet skills:
- ai-engineering
- qa-engineering memory: user
You are a senior LLM evaluat
Alternatives in Data & AI
- Claude Mem — A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresse 38.9k ★
- Docs Drift Reviewer — Checks whether clickhouse-go documentation matches changes to its ClickHouse API, database/sql API, transports 3.3k ★
- Patina Naturalness Reviewer — Triggers to re-scan a patina rewrite for residual AI tells and over-editing risk 352 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Evals Engineer
Use for anything under evals/ — authoring or curating golden-set tasks, implementing the outcome/trajectory/co
AI Risk Officer
Owns AI-specific risk in both directions — AI features the team ships, and the AI agents the team builds with.
Wikiskill Grader
Strict LLM-judge grader — given a task's success criteria and an inference agent's transcript/final answer, re
Mx Shipkit Builder
Use this agent to integrate the Ship Kit (analytics, SEO, Stripe payment, feedback widget, contact footer) int
Brand Equity Health Tracking Subagent
Sub-agent owning Brand Equity Measurement & Brand Health Tracking — designing the measurement framework (aware
Solidjs Architect
Use for SolidJS / SolidStart architectural work — scaffolding a new app, organizing folders, deciding where st
Related Skills
Agent Eval Workbench
MLflow-based evaluation kit for AI agents: golden-task runner, calibrated LLM judge, list-price cost accountin
Wireframe Doc
Figma is for designing; this is for deciding what to design. Low-fi product wireframes as one shareable URL —
Fde Simulation
Hands-on Forward Deployed Engineer role simulations: scope an enterprise engagement, build multi-agent prototy