Agent Eval Harness — AI skill for Claude Code
Live, open-source benchmark for comparing AI coding agents on real GitHub issues.
How to install Agent Eval Harness
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open linny006/agent-eval-harness and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Agent Eval Harness does
Live, open-source benchmark for comparing AI coding agents on real GitHub issues.
Alternatives in AI
- Kilocode — GitHub Repo stars Open Source AI coding assistant for planning, building, and fixing code Open source AI assis 17k ★
- Osaurus — Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cry 5.1k ★
- CodeIsland — Real-time AI coding agent status panel in your MacBook notch — live status, approvals & replies for 13 AI tool 2.3k ★
README
Agent Eval Harness
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
[](https://github.com/linny006-tecch/agent-eval-harness/stargazers) [](https://github.com/linny006-tecch/agent-eval-harness/commits) [](#) [](#)
**⭐ Star this repo to bookmark — fresh data every 15 minutes**
[English](./README.md) · [中文](./README_CN.md) · [日本語](./README_JA.md) · [한국어](./README_KO.md) · [Español](./README_ES.md) · [Português](./README_PT.md)
💡 What is this?
A standardized benchmark suite that runs coding agents against live, real-world GitHub issues with reproduction steps. Unlike static academic benchmarks, it outputs a weekly-updated public leaderboard, enabling developers to compare agents like OpenCode, Codex, and Claude Code in realistic scenarios.
This list is **auto-updated every 15 minutes** by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.
📋 Current Items
⏰ Last updated: 2026-09-12 01:31 UTC
Data source: `GitHub Search API`
The table below is rewritten on every cron tick. Star the repo to bookmark.
| # | Name | ⭐ | Lang | Updated | Description |
|---|---|---|---|---|---|
| 1 | promptfoo/promptfoo | 25037 | TypeScript | 2026-09-12 | Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare perform |
Related Skills
PIBench
PIBench is an open-source benchmark for evaluating AI coding agents on realistic, end-to-end payment integrati
BossConsole
Open-source, multi-platform harness for AI agents — a native, multi-threaded operator's console (JVM, not Elec
Harness Arena
Blind, LMSYS-style arena for comparing AI coding agent harnesses (Claude Code, Codex CLI, OpenClaw, Hermes, On
Harness Arena Fe
Blind, LMSYS-style arena for comparing AI coding agent harnesses (Claude Code, Codex CLI, OpenClaw, Hermes, On
Aura Code
Model-agnostic autonomous coding agent with a reproducible benchmark harness, adaptive turn budgets, and domai
AI Bridge
Local bridge letting an AI coding agent drive your real, logged-in Chrome — trusted input, CSP-proof eval, aut
Related Agents
Hcs Eval Reviewer
Reviews regression-trap quality and eval harness coverage for HCS. Ensures traps capture real failure classes
Harness Code Reviewer
Code reviewer — two-stage review against a pinned SHA: spec compliance first, then code quality, hunting fail-
🦀 Claw CR
Der Open-Source-Coding-Agent, in Rust gebaut. Alternative zu ClaudeCode Status Language