Cc Judge — AI skill for Claude Code
Shared LLM-evals framework for running Claude Code agents in containers, scoring outcomes with an LLM judge, and producing reports.
How to install Cc Judge
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open chughtapan/cc-judge and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Cc Judge does
Shared LLM-evals framework for running Claude Code agents in containers, scoring outcomes with an LLM judge, and producing reports. Extracted from moltzap's packages/evals so any project can depend on it.
Alternatives in AI
- System Prompts Leaks — Extracted system prompts from ChatGPT (GPT-5.4, GPT-5.3, Codex), Claude (Opus 4.6, Sonnet 4.6, Claude Code), G 38.6k ★
- Codex Plugin — 's /codex:adversarial-review command 9.6k ★
- Contributing To Claude Skills Library — Thank you for your interest in contributing to the Claude Skills Library 5.3k ★
README
cc-judge
TypeScript-first CLI + SDK for planned Claude Code harness runs, LLM bundle judging, and `summary.md` + `results.jsonl` reports. Telemetry can fan out to Braintrust and Promptfoo through pluggable emitters.
Prerequisites
- Node.js 20.11+
pnpm9+- Claude judge auth through
claude auth loginorANTHROPIC_API_KEY=... - Docker, if your harness launches Docker workloads
When `--judge-backend anthropic` is active and `ANTHROPIC_API_KEY` is not set, `cc-judge` runs `claude auth status` before `run`. Successful checks are cached for 24 hours under the user cache directory.
Install
pnpm add cc-judge
Quickstart
Create a harness-backed plan:
project: moltzap
scenarioId: EVAL-005
name: Cold outreach response quality
description: Verify the target agent responds helpfully to a first-contact DM.
requirements:
expectedBehavior: The agent should answer coherently instead of returning an auth or runtime error.
validationChecks:
- Response contains non-empty text
- Response stays on topic
harness:
module: ../../packages/runtimes/dist/trace-capture-harness.js
payload:
runtime:
kind: openclaw
conversation:
kind: direct
setupMessage: Hello, can you explain how MoltZap conversations work?
Run it:
cc-judge run ./plans/**/*.yaml --results ./eval-results --log-level info
Successful and failed runs both emit:
eval-results/
summary.md
results.jsonl
details/
..yaml
Harness Modules
The plan's `harness.module` path resolves relative to the plan file. The module must export `load(args)` as its default export unless the plan sets `harness.export`.
Minimal module shape:
import { Effect } from "effect";
export default {
load(args) {
return Effect.succeed({
plan: {
project: args.plan.project,
scenarioId: args.plan.scenarioId,
name: args.plan.name,
description: args.plan.descripti
Related Skills
Ground Truth Evals
An LLM eval harness graded by computed ground truth, not an LLM judge. Worked example: poker, served to models
Team Foundry
One shared context layer for Claude Code, Cursor, Gemini CLI, and Codex — so the whole team's AI tools build f
Geo SEO Claude
GEO-first SEO skill for Claude Code. Comprehensive AI search optimization for any website — citability scoring
Team Memory MCP
Shared team memory for AI coding agents. Bayesian confidence scoring with temporal decay. Works with Claude Co
Full Stack Chat Application With Multi Platform Clients AI Agent Integration
A Claude Code-style AI agent harness in Python. Uses Ollama + native LLM tool-calling to read/write files and
Aqven
Python framework + local Studio for reliable LLM workflows: typed files in your repo, pre-run checks, experime
Related Agents
Evals Engineer
Use for anything under evals/ — authoring or curating golden-set tasks, implementing the outcome/trajectory/co
DevOps Agent
DevOps teammate — infrastructure, CI/CD, containers, configuration, and documentation. Claims tasks from the s
AI Data Specialist
Deep AI/data engineer — LLM integration and agent systems, RAG, evals, data pipelines, and ML productionizatio