Aacr Bench
Description
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Installation
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open the source below and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
README

[](LICENSE) [](https://arxiv.org/abs/2601.19494) [](https://huggingface.co/datasets/Alibaba-Aone/aacr-bench)   
English | [简体中文](README.zh-CN.md)
📋 Introduction
AACR-Bench is the **industry's first multilingual, repository-level context-aware code review evaluation dataset**, designed to assess the performance of large language models in automated code review tasks. The dataset comprises 200 real Pull Requests from 50 active open-source projects, covering 10 mainstream programming languages. Each instance includes not only code changes but also preserves complete repository context, authentically reproducing the entire code review process. Through human-LLM collaborative review combined with multi-round expert annotation, we ensure high quality and comprehensiveness of the data. 
✨ Core Features
🌍 **Multi-language Coverage**
Covers **10** mainstream programming languages used in projects:
- System-level languages: C++, Rust, Go
- Enterprise languages: Java, C#, TypeScript
- Scripting languages: Python, JavaScript, Ruby, PHP
📁 **Repository-level Context**
- Preserves complete project structure
- Supports cross-file references and inter-module interaction analysis
- Includes PR metadata (description, title, comments, etc.)
🤖 **Human Expert + LLM Enhanced Annotation**
**Professional Annotation Team**
- 80+ senior software engineers with 2+ years of experience
Related Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI TrendRadar
⭐AI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.🎯 告别信息过载,你的
AI mem0
| Universal memory layer for AI Agents | 51341 | 221 | 1 |
AI