Blind Evaluator — Development agent for Claude Code
Structurally separate eval agent.
How to install Blind Evaluator
Installs to ~/.claude/agents/ariaxhan-kernel-claude-blind-evaluator.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/ariaxhan/kernel-claude/HEAD/agents/blind-evaluator.md -o ~/.claude/agents/ariaxhan-kernel-claude-blind-evaluator.md Restart Claude Code, or start a new session, for it to be picked up.
What Blind Evaluator does
name: blind-evaluator description: "Structurally separate eval agent. Receives ONLY the problem statement + rubric, NEVER the solution or the implementing agent's output. Used for high-stakes assessment where self-scoring would inflate the result." tools: Read, Bash, Grep, Glob
The eval agent that doesn't know the answer. You receive: problem statement, success rubric, optional reference behavior. You DO NOT receive: the implementing agent's output,Alternatives in Development
- Engram Assessor — Independent grader of learner productions for the Engram learning plugin 1.4k ★
- Synthesis Composer — Optional decision helper that turns Vestige retrievals into concise recommendations 632 ★
- Python Sidecar Specialist — Use for the Python sidecars on both surfaces: the verify checks (syntax, imports, statement mapping, coverage 179 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Legacy State Auditor
Version adversary - finds where current code mishandles persisted state that only a PAST version (or external
Bench Reviewer
Reviews eval suite quality — fixture/rubric consistency, scoring accuracy, false positive traps, difficulty ca
Fixture Builder
Creates and enriches eval suite fixtures — .md component files, .metadata.yaml grading criteria, and .rubric.y
Epo Patent Drafter
Expert in drafting EPO-compliant patent claims and specifications. Specializes in two-part form claims, proble
Product Market Fit Analyst
Product-market fit specialist for validating problem-solution alignment and market readiness
Claudehut Brainstormer
Generates 2-4 genuinely distinct solution options scored against locked criteria and recommends one. Any probl
Related Skills
Blind
Strategy-level blind drill on a known HW or example problem. User describes approach in prose (no math typing)
Rubric Evaluator
Evaluate Claude Code and Codex skill directories with deterministic rubric checks and graded, fixable reports.
Evaluator Library
Dispatch a fresh-context evaluator against a domain rubric. Bundled domains: code-quality, ui-design, prose, s