Sme Eval Triage — Development agent for Claude Code
Triage a golden-set failure from the compliance-SME seat before anyone edits ground truth.
How to install Sme Eval Triage
Installs to ~/.claude/agents/andaro74-regdelta-sme-eval-triage.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/andaro74/regdelta/HEAD/.claude/agents/sme-eval-triage.md -o ~/.claude/agents/andaro74-regdelta-sme-eval-triage.md Restart Claude Code, or start a new session, for it to be picked up.
What Sme Eval Triage does
name: sme-eval-triage description: Triage a golden-set failure from the compliance-SME seat before anyone edits ground truth. Use whenever `make evals` fails or a golden-question change is proposed. tools: Read, Grep, Glob, WebSearch, WebFetch
You are the regulatory-affairs SME for RegDelta. A human SME owns ground truth; you prepare their decision. For each failing question:
- Classify — exactly one: a. MODEL/SYSTEM REGRESSION — the regulation is unchanged; the system's answe
Alternatives in Development
- Review Agent — You are a specialized review agent 3.6k ★
- Eval Engineer — GAIA evaluation framework specialist 1.5k ★
- Packager Troubleshooter — Use this agent when packaging fails or needs changes — yarn build:package:windows mac linux, PyInstaller error 204 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Bizreq Analyst
Business-requirements agent — authors business-rules.md (a BR-NN decision table) and golden-scenarios.md (SCEN
Sonmat Witness
External witness agent. Verifies intent-artifact match using user turn cascade and ground truth. Protocol-isol
Stroi Explorer
Use this agent when a planner or skill needs ground truth about the codebase before designing against it. Typi
Doc Analyser
Per-page documentation analyser for the ground-truth doc-lineage layer. Reads one published doc page (the GitB
Cutter
Turn recorded biographies into the shaped scene-list of a novel — the documentary editor's cut. Given the exha
X402 Economy Triage
Diagnoses x402 economy outages (settle failures, dry wallets, fee-wallet floor starvation) using the known fai
Related Skills
Test Upload Evals Skill
End-to-end test of publishing a skill with evals on the dev environment. Run each step sequentially — stop and
Rails AI Context
45 MCP tools that give AI coding agents ground truth about your Rails app: schema, models, routes, controllers
Claude Boot Stats
Per-component token cost analyzer for Claude Code first-turn context: HTTP intercept proxy + Qwen3 BPE + groun