Bench Watch
Description
--- description: Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart per-arm-model summary (catch-rate AND cry-wolf). For the run-slice / wait-for-it / show-the-per-arm-model-summary loop. allowed-tools: Bash, Read, Glob, Grep, Task --- ## Context Plumbline's identity is *measure, don't assert*: every claim earns a deterministic bench. Bench runs are long, and the manual loop "wait for the slice to finish → show the per-arm-model summary"
Installation
Installs to ~/.claude/commands/dyai2025-plumbline-bench-watch.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/DYAI2025/Plumbline/HEAD/.claude/commands/bench-watch.md -o ~/.claude/commands/dyai2025-plumbline-bench-watch.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Awesome Llm Apps
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI