Regression Sentinel — Development agent for Claude Code
Watch evaluation metrics over time for trends and regressions.
How to install Regression Sentinel
Installs to ~/.claude/agents/kastalien-research-thoughtbox-regression-sentinel.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/Kastalien-Research/thoughtbox/HEAD/.claude/agents/regression-sentinel.md -o ~/.claude/agents/kastalien-research-thoughtbox-regression-sentinel.md Restart Claude Code, or start a new session, for it to be picked up.
What Regression Sentinel does
name: regression-sentinel description: Watch evaluation metrics over time for trends and regressions. Unlike verification-judge (validates single work items against specs), this agent watches metric trends across sessions and flags gradual degradation before it becomes critical. Use after session-end metrics collection or as a daily rollup. tools: Read, Glob, Grep, Bash, ToolSearch disallowedTools: Edit, Write model: sonnet maxTurns: 12 memory: project
You are the Regression Sentinel Ag
Alternatives in Development
- Gitnexus PR Facts Historian — GitNexus PR facts and repository-history investigator 45.8k ★
- Prometheus Configuration — Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules 31.9k ★
- Civitai Perf Review — Reviews a feature segment in the main Civitai Next.js app (src/) for production performance — N+1 queries, uni 7.2k ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
PR Check Watcher
Watches a GitHub PR's CI checks to completion via gh pr checks --watch and reports a concise pass/fail summary
Metric Optimizer
Bounded metric-ratchet specialist for direction-not-destination work — proposes a change, applies it, runs the
Doey Critic
Quality critic — fast, ruthless, minimal. Reviews code and output for correctness, clarity, and necessity. Own
Legacy Modernizer
Refactor legacy codebases, migrate outdated frameworks, and implement gradual modernization. Handles technical
Prose Polisher
Rewrites existing text to improve clarity, conciseness, flow, and adherence to academic writing principles. Un
Refactor Specialist
Use this agent for behavior-preserving refactors — restructuring, renaming, extracting, de-duplicating, and un
Related Skills
Rogue ML Failure Audit Skill
General workflow for auditing ML CI failures, experiment regressions, training run failures, golden metric fai
Update Rules
Interactively propose ALLOW/ASK rule additions for agent-sentinel from recent LLM_JUDGE log entries
Secnews
面向 AI + 安全从业者的单人本地工作站: Sentinel Terminal 三层工作流 (Data 采集 / Judge 评估 / Action 执行) · llm-wiki-2.0 wiki-first 知识库