HomericIntelligence

Benchmark Specialist — Data & AI agent for Claude Code

Data & AI community

Use for benchmark execution, monitoring, and data collection during evaluation runs.

How to install Benchmark Specialist

Installs to ~/.claude/agents/homericintelligence-scylla-benchmark-specialist.md

Terminal
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/HomericIntelligence/Scylla/HEAD/.claude/agents/benchmark-specialist.md -o ~/.claude/agents/homericintelligence-scylla-benchmark-specialist.md

Restart Claude Code, or start a new session, for it to be picked up.

What Benchmark Specialist does


name: benchmark-specialist description: Use for benchmark execution, monitoring, and data collection during evaluation runs. Invoked for running tier benchmarks and collecting raw evaluation data. tools: Read,Write,Edit,Bash,Grep,Glob model: sonnet

Benchmark Specialist Agent

Role

Level 3 Specialist responsible for executing benchmarks and collecting evaluation data. Manages benchmark runs across tiers, monitors execution, and ensures data quality.

Hierarchy Position

  • **Leve

Alternatives in Data & AI

  • Callstackincubator — A collection of agent-optimized React Native skills for AI coding assistants 1k ★
  • Mathodology Evidence Researcher — Use for literature, data source, background, benchmark, and citation work in award-level modeling submissions 137 ★
  • Peer Reviewer Ethics — A simulated peer reviewer specializing in research ethics, IRB compliance, informed consent, vulnerable popula 134 ★

Full documentation available on GitHub

View Source Repository