linny006

Agent Eval Harness — AI skill for Claude Code

AI community

Live, open-source benchmark for comparing AI coding agents on real GitHub issues.

How to install Agent Eval Harness

This entry records only its repository, not the path inside it, so there is no exact command to give. Open linny006/agent-eval-harness and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Agent Eval Harness does

Live, open-source benchmark for comparing AI coding agents on real GitHub issues.

Alternatives in AI

  • Kilocode — GitHub Repo stars Open Source AI coding assistant for planning, building, and fixing code Open source AI assis 17k ★
  • Osaurus — Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cry 5.1k ★
  • CodeIsland — Real-time AI coding agent status panel in your MacBook notch — live status, approvals & replies for 13 AI tool 2.3k ★

README

Agent Eval Harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

[![Stars](https://img.shields.io/github/stars/linny006-tecch/agent-eval-harness?style=for-the-badge&logo=github)](https://github.com/linny006-tecch/agent-eval-harness/stargazers) [![Last Commit](https://img.shields.io/github/last-commit/linny006-tecch/agent-eval-harness?style=for-the-badge)](https://github.com/linny006-tecch/agent-eval-harness/commits) [![Items](https://img.shields.io/badge/Tracked_Items-100-brightgreen?style=for-the-badge)](#) [![Updated](https://img.shields.io/badge/Updates-every_15min-blue?style=for-the-badge)](#)

**⭐ Star this repo to bookmark — fresh data every 15 minutes**

[English](./README.md) · [中文](./README_CN.md) · [日本語](./README_JA.md) · [한국어](./README_KO.md) · [Español](./README_ES.md) · [Português](./README_PT.md)


💡 What is this?

A standardized benchmark suite that runs coding agents against live, real-world GitHub issues with reproduction steps. Unlike static academic benchmarks, it outputs a weekly-updated public leaderboard, enabling developers to compare agents like OpenCode, Codex, and Claude Code in realistic scenarios.

This list is **auto-updated every 15 minutes** by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.


📋 Current Items

⏰ Last updated: 2026-09-12 01:31 UTC

Data source: `GitHub Search API`

The table below is rewritten on every cron tick. Star the repo to bookmark.

# Name Lang Updated Description
1 promptfoo/promptfoo 25037 TypeScript 2026-09-12 Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare perform