ResearchClawBench banner
InternScience InternScience

ResearchClawBench

AI community

Description

🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

ResearchClawBench

[![Official Site](https://img.shields.io/badge/Official%20Site-333399.svg?logo=homepage)](https://InternScience.github.io/ResearchClawBench-Home/)  [![GitHub](https://img.shields.io/badge/GitHub-000000?logo=github&logoColor=white)](https://github.com/InternScience/ResearchClawBench)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-gray)](https://huggingface.co/datasets/InternScience/ResearchClawBench)  [![ModelScope](https://img.shields.io/badge/%F0%9F%A4%96%20ModelScope-9B00FF.svg)](https://www.modelscope.cn/datasets/InternScience/ResearchClawBench)  [![arXiv](https://img.shields.io/badge/arXiv-paper-b31b1b.svg?logo=arxiv&logoColor=white)](https://arxiv.org/pdf/2606.07591)  [![Domains](https://img.shields.io/badge/Domains-10-green.svg)](#-scientific-domains) [![Tasks](https://img.shields.io/badge/Tasks-40-orange.svg)](#-scientific-domains) [![GitHub](https://img.shields.io/github/stars/InternScience/ResearchClawBench?style=social)](https://github.com/InternScience/ResearchClawBench)

**Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery**

[Quick Start](#-quick-start) | [Submit Tasks](#-submit-new-tasks) | [How It Works](#%EF%B8%8F-how-it-works) | [Domains](#-scientific-domains) | [Leaderboard](#-leaderboard) | [Add Your Agent](#-add-your-own-agent)

SGI Overview


ResearchClawBench is a benchmark that measures whether AI coding agents can **independently conduct scientific research** β€” from reading raw data to producing publication-quality reports β€” and then rigorously evaluates the results against **real human-authored papers**.

Unlike benchmarks that test coding ability or factual recall, ResearchClawBench asks: *given a curated scientific workspace and the same research goal, can an AI agent arrive at the same (or better