Agent Eval Hub — AI skill for Claude Code
Multi-provider + cross-device reliability harness for AI agents.
How to install Agent Eval Hub
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open dhavig/agent-eval-hub and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Agent Eval Hub does
Multi-provider + cross-device reliability harness for AI agents. Runs the same suite across Claude, OpenAI, Gemini, Ollama, and Android surfaces - scoring task success, tool calls, safety, cost, latency, and drift with statistical A/B testing
Alternatives in AI
- Artifacts Builder — Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech 97.5k ★
- J Space Cognition Suite V3.7 — J-Space Cognition Suite V3.7 - AI cognitive-enhancement Skills based on Anthropic's J-space global workspace r 3k ★
- Loli Profiler — Memory instrumentation tool with CI and AI support for android app&game developers 695 ★
README
AgentEvalHub
Multi-provider + cross-device reliability harness for AI agents. Runs the same agent spec against Claude / OpenAI / Gemini / Ollama and against real or mocked Android surfaces. Scores every run on task success, tool-call correctness, safety under attack, cost, latency, and regression over time.
Built as a QA portfolio project for the agentic-AI era.
Quick start
git clone
cd agent-eval-hub
python -m venv .venv && source .venv/bin/activate
pip install -e .[dev]
cp .env.example .env
# fill in ANTHROPIC_API_KEY, OPENAI_API_KEY, and/or GEMINI_API_KEY
export $(cat .env | xargs)
# Run the same suite against two providers
agent-eval --suite suites/agent/tool_use.yaml --provider claude --model claude-sonnet-4-6
agent-eval --suite suites/agent/tool_use.yaml --provider openai --model gpt-4o-mini
# Compare two surfaces for answer agreement
agent-eval-cross --suite suites/agent/tool_use.yaml \
--surface-a claude:claude-sonnet-4-6 --surface-b ollama:llama3.1
# A/B test with statistical significance
agent-eval-ab --suite suites/safety/red_team.yaml \
--surface-a claude:claude-sonnet-4-6 --surface-b claude:claude-haiku-4-5-20251001
# Cross-surface safety parity (refusal behavior must match)
agent-eval-safety-parity --suite suites/safety/red_team.yaml \
--surface-a claude:claude-sonnet-4-6 --surface-b device:llama3.1
# Dashboard + synthetic demo data
agent-eval-seed --db demo.duckdb
streamlit run src/agent_eval_hub/dashboard/app.py -- --db demo.duckdb
# Tests
pytest # everything local (93+ tests)
pytest -m "not integration and not e2e" # fast unit tier
Modules shipped
| Module | What |
|---|---|
| 1 | Provider-agnostic agent loop (Claude / OpenAI / Gemini / Ollama) |
| 2 | RAG grounding suite + LLM-as-judge with defensive JSON parsing |
| 3 | Red-team suite (5 attack classes) + safety graders (refused, did_not_contain, did_not_call_tool) |
Related Skills
LLM Dark Patterns
Umbrella for the LLM Dark Patterns Hooks suite — single-purpose Claude Code Stop hooks that suppress sycophanc
Shellbench
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based sc
Wishing Willow
Show what the model thinks you asked, next to what you actually said. A Claude Code plugin that surfaces task-
Chat Hub
Desktop cockpit for AI coding agents: Claude Code / Codex / Grok / OpenCode side by side — worktree-per-sessio
Driftproof
Continuous, model-version-bound verification of agent skills: run a skill's eval suite with and without the sk
AI Gate
AI verification for G4 — eval suite against baseline, red-team status, guardrail verification, drift check
Related Agents
ZeroClaw Android
Run AI agents 24/7 on your Android phone. Native Rust core, 25+ providers (OpenAI, Claude, Gemini, Groq, DeepS
AI ML
AI/ML 통합 전문가 + LLM API 최신 모델/SDK 코딩 가이드. RAG 시스템, 문서 분석, OpenAI/Anthropic/Gemini/Ollama 최신 API 보장. "AI integra
LLM Orchestrator
Use this agent for LLM integration work — prompt engineering, multi-provider abstraction (OpenRouter, Gemini,