Gov Eval — Testing skill for Claude Code
Run GovEval — the governance validation test suite.
How to install Gov Eval
Installs to ~/.claude/commands/datallmhub-claude-governance-gov-eval.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/datallmhub/claude-governance/HEAD/.claude/commands/gov-eval.md -o ~/.claude/commands/datallmhub-claude-governance-gov-eval.md Restart Claude Code, or start a new session, for it to be picked up.
What Gov Eval does
Run GovEval — the governance validation test suite.
What this command does
Runs automated tests that verify the governance rules are actually followed. Each scenario sends a natural prompt to the generator, then an independent judge model scores the output.
Prerequisites
MISTRAL_API_KEYenvironment variable must be set (judge model)- Python 3.11+ with dependencies installed:
cd /tests && pip install -r requirements.txt
Steps
- Read `.claude-governanc
Alternatives in Testing
- Test Full — Run the full test suite across all tox environments with coverage enabled 1.1k ★
- Check Package — Run full validation (build + lint + test + typecheck) on a package 1.1k ★
- Test Marker — Create a simple test marker file for POC validation 229 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Eval Suite
Re-score tiny-spec against the SDD evaluation rubric, record it, and report the delta vs the last run.
Integration Test Command
Execute comprehensive integration test suite with multi-component coordination and validation.
Fix And Validate
Run the Lancet2 fix-mode validation suite (iwyu-fix + lint-check + test). Mirrors the pre-commit gate's scope
Gen Evals
Generate EVAL-.md test cases for an agent from its prompt. Usage: /gen-evals [--count N]
Skill Eval Action
GitHub Action to evaluate Claude Code skills against YAML test cases with automated grading and PR reporting
Agent Eval Design
Design an evaluation for an AI agent or LLM feature: what to test, how to grade it, and how to catch regressio
Related Agents
Deployment Cicd Analyst
Generic deployment and CI/CD governance analyst for BC Gov engagements. Produces evidence-based findings, cont
Prosa Test Runner
Runs the prosa validation suite and reports failures concisely. Use to validate a change end-to-end or refresh
Test Lead
Use for test strategy, test suite design, and QA validation of completed work. Designs tests, delegates execut