Agentgovbench — Testing skill for Claude Code
48-scenario benchmark testing identity, policy enforcement, and observability across AI agent runtimes.
How to install Agentgovbench
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open agentic-control-plane/agentgovbench and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Agentgovbench does
48-scenario benchmark testing identity, policy enforcement, and observability across AI agent runtimes.
Alternatives in Testing
- Webapp Testing — Test local web applications using Playwright for UI verification and debugging 94.1k ★
- Fix Issue — by metabase - Addresses GitHub issues by taking issue number as parameter, analyzing context, implementing sol 46.5k ★
- Playwright — claude-plugins-official Browser automation, E2E testing, screenshots 29.4k ★
README
AgentGovBench
An open benchmark for AI agent governance. Mapped to NIST AI RMF. Vendor-neutral.
Live scorecard · All 48 scenarios · Methodology · Architecture-is-governance · agenticcontrolplane.com
What it measures
Existing benchmarks (HarmBench, InjecAgent, AgentDAM, AgentLeak) test the **model** — does the LLM refuse harmful prompts, resist injection, protect PII. AgentGovBench tests the **governance layer around the model** — the part responsible for who can call which tool, whose identity rides along with each call, how rate limits cascade across delegated subagents, and what the audit record contains after the fact.
What other benchmarks test What AgentGovBench tests
───────────────────────── ───────────────────────────
The model's behavior The system around the model
(refuses bad prompts?) (enforces the policy?)
(attributes the call
Related Skills
Ultimate Product Discovery Skill
Ultimate Claude Skill for Product Discovery. 18 tasks across 6 blocks: Market Analysis, Customers (JTBD/CJM),
Ralph Output Adversarial Review
Adversarially reviews the artifacts produced by a RALPH run (code/tests/UI/config) against repo policy (testin
Cceval
Evaluate and benchmark your CLAUDE.md effectiveness with automated testing
Decision Quality Analyzer
Analyze team decision quality with bias detection, scenario testing, and process improvement recommendations
Council Review
Multi-perspective senior review board. Debate across architecture, UX, implementation risk, and testing. Conve
Cc Codex Triage
Claude Code plugin for persistent triage dialogue with the OpenAI Codex CLI via codex exec resume — cross-agen
Related Agents
Testing QA
🧪 Use this agent for test design and execution, bug hunting and reproduction, edge-case and coverage analysis
Audit Policy Compliance
Platform policy specialist. Returns schema-valid findings covering platform eligibility, regulated categories,
Dotnet Testing Specialist
WHEN designing test architecture, choosing test types (unit/integration/E2E), managing test data, testing micr