AI Agent Eval Writer — Testing & QA agent for Claude Code
Use this agent when you need to write evaluation tests for AI agents using Evalite or Autoevals frameworks.
How to install AI Agent Eval Writer
Installs to ~/.claude/agents/januarylabs-deepagents-ai-agent-eval-writer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/JanuaryLabs/deepagents/HEAD/.claude/agents/ai-agent-eval-writer.md -o ~/.claude/agents/januarylabs-deepagents-ai-agent-eval-writer.md Restart Claude Code, or start a new session, for it to be picked up.
What AI Agent Eval Writer does
name: ai-agent-eval-writer description: Use this agent when you need to write evaluation tests for AI agents using Evalite or Autoevals frameworks. This includes creating evaluation suites to test LLM outputs, agent behaviors, response quality, factual accuracy, or any other AI system performance metrics. Examples of when to invoke this agent:\n\n\nContext: The user has just finished implementing an AI agent or LLM-based feature and wants to ensure it performs correctly.\nuser: "I j
Alternatives in Testing & QA
- Rivetkit Driver Test Writer — Use this agent when the user needs to write, modify, or debug driver tests for RivetKit 6.1k ★
- Testing Expert — Use this agent for Output.ai testing strategies including Vitest configuration, Temporal workflow testing, LLM 434 ★
- Module Test Writer — Use this agent to write feature tests for LaraDashboard module CRUD operations, services, policies, and Livewi 407 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Logic Architect
Extract capabilities and behaviors from spec. Outputs JSON that MUST validate against schemas/capability-map.s
AI Prompt Architect
Designs and versions LLM system prompts for ai-system / agent-product archetypes. Outputs docs/decisions/ADR-{
Airlock Test Writer
Write Airlock tests with Test::More. Never run the live suites — they hit a real Keycloak and are opt-in only.
Provider Tester
Use this agent when you need to test, debug, or validate LLM provider configurations. This includes verifying
Eval Auditor
Audits an evaluation setup (benchmark, A/B test, or model comparison) for methodology errors that would invali
AI Eval Designer
Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design ev
Related Skills
Audit E2e
Audit end-to-end test coverage of critical Playwright paths. Identifies missing flows, flaky tests, slow suite
Ashlr Eco Mode
Toggle eco mode — aggressive token-saving behaviors that trade some response richness for lower session cost.
Mini Claude Mods
Mini mods for Claude Code: function-hook plugins (panes, commands, behaviors) plus a marketplace to install th