Evaluation Orchestrator — Development agent for Claude Code
Use for coordinating evaluation experiments, managing experiment execution, and overseeing evaluation workflows.
How to install Evaluation Orchestrator
Installs to ~/.claude/agents/homericintelligence-scylla-evaluation-orchestrator.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/HomericIntelligence/Scylla/HEAD/.claude/agents/evaluation-orchestrator.md -o ~/.claude/agents/homericintelligence-scylla-evaluation-orchestrator.md Restart Claude Code, or start a new session, for it to be picked up.
What Evaluation Orchestrator does
name: evaluation-orchestrator description: Use for coordinating evaluation experiments, managing experiment execution, and overseeing evaluation workflows. Invoked for running tier comparisons and evaluation studies. tools: Read,Write,Edit,Bash,Grep,Glob,Task model: sonnet
Evaluation Orchestrator Agent
Role
Level 1 Domain Orchestrator responsible for coordinating evaluation experiments. Manages the execution of tier comparisons, oversees data collection, and ensures experiment in
Alternatives in Development
- Sales Automator — Draft cold emails, follow-ups, and proposal templates 2.8k ★
- Gpui Researcher — Researches and validates GPUI usage patterns, APIs, and conventions 2.2k ★
- Eval Engineer — GAIA evaluation framework specialist 1.5k ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Autoresearch Orchestrator
Use this agent to run an autoresearch experiment session end-to-end — setup, baseline, and a batch of edit→mea
Experiment
Performs focused technical experiments for uncertain behavior, algorithms, geometry, rendering, or performance
Baime Iteration Executor
Executes a single experiment iteration through its lifecycle phases. This involves coordinating Meta-Agent cap
Biz Release Manager
Release manager who coordinates deployments, manages release schedules, and oversees go-live processes. Use th
Skill Eval Reporter
Compares repeated paired execution results using blind A/B methodology and generates a skill effectiveness rep
Bully Evaluator
Evaluates a single bully semantic-evaluation payload against a diff and returns a structured violation list. I
Related Skills
Experiment Design
Design empirical studies through power analysis, pre-analysis planning,
Clone Study
A Claude Code skill that studies and recreates high-end websites (Awwwards-tier) from a URL — capture, analyze
Claude Hook Experiment
Reproducible Claude Code hook experiments for validating settings-based hook behavior