Cceval — Testing skill for Claude Code
Evaluate and benchmark your CLAUDE.md effectiveness with automated testing.
How to install Cceval
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open johnlindquist/cceval and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Cceval does
Evaluate and benchmark your CLAUDE.md effectiveness with automated testing.
Alternatives in Testing
- Webapp Testing — Test local web applications using Playwright for UI verification and debugging 94.1k ★
- Fix Issue — by metabase - Addresses GitHub issues by taking issue number as parameter, analyzing context, implementing sol 46.5k ★
- Playwright — claude-plugins-official Browser automation, E2E testing, screenshots 29.4k ★
README
cceval
Evaluate and benchmark your CLAUDE.md effectiveness with automated testing.
Why?
Your `CLAUDE.md` file guides how Claude Code behaves in your project. But how do you know if your instructions are actually working?
**cceval** lets you:
- Test different system prompt variations
- Compare performance across metrics
- Find what actually improves Claude's behavior
- Share benchmarks with your team
Installation
# Global install (recommended)
bun add -g cceval
# Or per-project
bun add -D cceval
**Requirements:**
- Bun ≥ 1.0
- Claude Code CLI installed and authenticated
Quick Start
# Run with default prompts and variations
cceval run
# Generate report from existing results
cceval report evaluation-results.json
# Create custom config
cceval init
Usage
Basic Evaluation
Run with sensible defaults (5 variations × 5 prompts = 25 tests):
cceval run
This tests:
- baseline: Minimal "You are a helpful assistant"
- gateFocused: Clear pass/fail criteria
- bunFocused: Bun-specific instructions
- rootCauseFocused: Root cause protocol
- antiPermission: Trust-based prompting
Against prompts that test:
- Reading files before coding
- Using Bun instead of Node
- Fixing root cause vs surface symptoms
- Asking permission appropriately
Custom Configuration
Create a config file:
cceval init
Edit `cceval.config.ts`:
import type { EvalConfig } from "cceval"
const config: EvalConfig = {
prompts: {
// Your test scenarios
authentication: "Add login functionality to the app.",
performance: "The dashboard is slow, optimize it.",
testing: "Add tests for the user service.",
},
variations: {
// Your system prompt variations
baseline: "You are a helpful coding assistant.",
myClaudeMd: `You are evaluated on gates. Fail any = FAIL.
1. Read files before coding
2. State plan then pr
Related Skills
Skill Eval Action
GitHub Action to evaluate Claude Code skills against YAML test cases with automated grading and PR reporting
Agentgovbench
48-scenario benchmark testing identity, policy enforcement, and observability across AI agent runtimes.
Vibetest Use
Vibetest MCP - automated QA testing using Browser-Use agents
ShipGuard
Claude Code skills for automated E2E testing with agent-browser (Playwright CLI). Discover routes, generate YA
AI Watch Tester
AI-powered automated E2E testing. Just enter a URL — AI generates and runs test scenarios.
Evolve Agent
Build agents that improve their own prompts through automated iteration and testing
Related Agents
Agent Prompt Reviewer
Use this agent when you need to evaluate and score Claude Code sub-agent prompts for quality and effectiveness
Learner Diagnostic
Assesses learning approach across 5 dimensions from Justin Sung's methodology. Use when user wants to evaluate
Evolve Benchmark Gate
Statistical benchmark comparison gate (Evaluate archetype).