Agent Eval — Productivity skill for Claude Code
Head-to-head comparison of coding agents / models on reproducible tasks, with pass-rate / cost / time / consistency metrics.
How to install Agent Eval
Installs to ~/.claude/skills/rustidi-ml-portfolio-agent-eval/SKILL.md
mkdir -p ~/.claude/skills/rustidi-ml-portfolio-agent-eval && curl -fsSL https://raw.githubusercontent.com/rustidi/ml-portfolio/HEAD/skills/agent-eval.md -o ~/.claude/skills/rustidi-ml-portfolio-agent-eval/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Agent Eval does
name: agent-eval description: Head-to-head comparison of coding agents / models on reproducible tasks, with pass-rate / cost / time / consistency metrics. Gives a framework instead of vibes when choosing between models for a step (e.g. transcription or text-polish). triggers: ["compare models", "which model is better", "agent eval", "head to head", "model comparison", "model selection"]
Agent Eval Skill
Sanitized example skill. Wraps a third-party CLI tool (credited below).
A ligh
Alternatives in Productivity
- /build-hook - Build Custom Claude Code Hooks — Interactive guide for creating custom hooks with Q&A workflow and automatic validation 618 ★
- Claude Code Usage Bar — Real‑time statusline for Claude Code: token usage, remaining budget, burn rate, and depletion time 159 ★
- Coder Eval Review — Generate per-task review.json (summary + tags) for a completed run 116 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Earshot
Blind, reproducible benchmark for production phone voice agents. Measures barge-in latency, false-stop rate an
Gov Metrics
Governance metrics — measure whether the gates actually work. Block rate, override/waiver rate, false-block pr
Agent Cleanfiles
Perform a general cleanup pass on the repository for periodic maintenance. For task-specific cleanup, use /age
Arena
Watch Browser Agents complete the same task head to head.
Cc Paw
CC Paw gives each project its own persistent Claude Code session with real-time status indicators and backgrou
Open Agent Octagon
Same-task comparison harness for coding agents (Claude Code, Codex, …). Not a leaderboard — it shows where and
Related Agents
Metrics Steward
Owns metric and KPI definitions, the metrics catalog, data dictionary, and measurement governance. Use when de
Chief Executive
Sets strategy and makes the final call on tradeoffs. Use when deciding what to build or not build, choosing be
Cost Control Reviewer
Reviews architecture and infrastructure for cost efficiency across AWS, Azure, and GCP. Use when provisioning