Bench Analyzer — Development agent for Claude Code
name: bench-analyzer description: Analyzes Lumen benchmark JSONL results to identify chunker/search quality issues and produce actionable improvement recommendations model: opus You are a benchmark analysis agent for….
How to install Bench Analyzer
Installs to ~/.claude/agents/ory-lumen-bench-analyzer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/ory/lumen/HEAD/.claude/agents/bench-analyzer.md -o ~/.claude/agents/ory-lumen-bench-analyzer.md Restart Claude Code, or start a new session, for it to be picked up.
What Bench Analyzer does
name: bench-analyzer description: Analyzes Lumen benchmark JSONL results to identify chunker/search quality issues and produce actionable improvement recommendations model: opus
You are a benchmark analysis agent for Lumen, a semantic code search tool. Your job is to analyze raw benchmark conversation logs, identify where Lumen's chunker and search failed, and produce actionable recommendations that generalize across codebases.
You will be given a benchmark results directory path.
Alternatives in Development
- Crash Analyzer Agent 2.3k ★
- Seam Analyzer — Hunts for a missing type at a seam — structure flattened and rebuilt downstream, hand-maintained lists held to 2.2k ★
- Issue Analyzer 2.1k ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Bench
API performance benchmarking — latency profiling, throughput testing, performance regression detection
Bench Reviewer
Reviews eval suite quality — fixture/rubric consistency, scoring accuracy, false positive traps, difficulty ca
Bench Runner
Executes a11y skill benchmarks across hosted and local model families. Runs cloud/Codex/Ollama benchmark scrip
Aurelius
Bench Crew coordinator. Fleet lead, canon author, Slack relay voice, morning digest drafter. Use for cross-tea
Bailey
Personal-space agent. Helps one authenticated user (by Bench UID) with their personal triage — inbox, follow-u
Process Expert
PhD process & automation engineer for copper/brass/cupronickel tube manufacturing. Use when question involves
Related Skills
People Search Bench
The first open benchmark for evaluating AI-powered people search agents
Bench
Run performance benchmarks for RLM-Claude-Code.
Bench Capture
Score the 2 bench variant folders for one size (build/acceptance/regression/conventions/capture + rubric + cro