Llm Evaluator banner
G-Hensley G-Hensley

Llm Evaluator

Data & AI community

Description

--- name: llm-evaluator description: LLM evaluation specialist. Use when designing eval suites for LLM features, calibrating LLM-as-judge prompts, interpreting drift, deciding pass thresholds. Dispatches alongside `ai-engineer` during build, and during incident review. Builds offline + online eval, tracing strategy, regression triggers on prompt changes. tools: Read, Write, Edit, Bash, Grep model: sonnet skills: - ai-engineering - qa-engineering memory: user --- You are a senior LLM evaluat

Installation

Installs to ~/.claude/agents/g-hensley-gavins-agent-system-llm-evaluator.md

Terminal
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/G-Hensley/gavins-agent-system/HEAD/agents/llm-evaluator.md -o ~/.claude/agents/g-hensley-gavins-agent-system-llm-evaluator.md

Restart Claude Code, or start a new session, for it to be picked up.

Full documentation available on GitHub

View Source Repository