Benchmark Specialist — Data & AI agent for Claude Code
Use for benchmark execution, monitoring, and data collection during evaluation runs.
How to install Benchmark Specialist
Installs to ~/.claude/agents/homericintelligence-scylla-benchmark-specialist.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/HomericIntelligence/Scylla/HEAD/.claude/agents/benchmark-specialist.md -o ~/.claude/agents/homericintelligence-scylla-benchmark-specialist.md Restart Claude Code, or start a new session, for it to be picked up.
What Benchmark Specialist does
name: benchmark-specialist description: Use for benchmark execution, monitoring, and data collection during evaluation runs. Invoked for running tier benchmarks and collecting raw evaluation data. tools: Read,Write,Edit,Bash,Grep,Glob model: sonnet
Benchmark Specialist Agent
Role
Level 3 Specialist responsible for executing benchmarks and collecting evaluation data. Manages benchmark runs across tiers, monitors execution, and ensures data quality.
Hierarchy Position
- **Leve
Alternatives in Data & AI
- Callstackincubator — A collection of agent-optimized React Native skills for AI coding assistants 1k ★
- Mathodology Evidence Researcher — Use for literature, data source, background, benchmark, and citation work in award-level modeling submissions 137 ★
- Peer Reviewer Ethics — A simulated peer reviewer specializing in research ethics, IRB compliance, informed consent, vulnerable popula 134 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Metrics Specialist
Use for metrics calculation, data collection, and measurement implementation. Invoked for calculating Pass-Rat
Executor Heavy
Orchestra heavy executor (Opus). Use for the hard tier of execution work orders — algorithmically hard cores (
Paper Section Drafter
Drafts a specific IMRAD section or paragraph for a scientific manuscript in the author's voice. Invoked by /pa
Data Executor
Internal dynos-work agent. Implements data pipelines, backfills, ETL, analytics, reconciliation, and data qual
Data Integrity Auditor
Internal dynos-work agent. Audits transactions, migrations, backfills, concurrency, idempotency, and data corr
LLM FinOps Architect
LLM FinOps Architect specializing in model-tier routing, cost-envelope governance, burn-rate monitoring, per-p
Related Skills
Fireworks Cost Benchmark
Public demo benchmarks of cost per successful execution
Gql Performance Benchmark
Run GraphQL performance benchmarks against the Alkemio API and compare against a stored baseline
Hatch3r Benchmark
Run and analyze performance benchmarks. Compare results against baselines, identify regressions, and produce p