ResearchClawBench
Description
π¦ ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Installation
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open the source below and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
README
ResearchClawBench
[](https://InternScience.github.io/ResearchClawBench-Home/) [](https://github.com/InternScience/ResearchClawBench) [](https://huggingface.co/datasets/InternScience/ResearchClawBench) [](https://www.modelscope.cn/datasets/InternScience/ResearchClawBench) [](https://arxiv.org/pdf/2606.07591) [](#-scientific-domains) [](#-scientific-domains) [](https://github.com/InternScience/ResearchClawBench)
**Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery**
[Quick Start](#-quick-start) | [Submit Tasks](#-submit-new-tasks) | [How It Works](#%EF%B8%8F-how-it-works) | [Domains](#-scientific-domains) | [Leaderboard](#-leaderboard) | [Add Your Agent](#-add-your-own-agent)
ResearchClawBench is a benchmark that measures whether AI coding agents can **independently conduct scientific research** β from reading raw data to producing publication-quality reports β and then rigorously evaluates the results against **real human-authored papers**.
Unlike benchmarks that test coding ability or factual recall, ResearchClawBench asks: *given a curated scientific workspace and the same research goal, can an AI agent arrive at the same (or better
Related Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Firecrawl
π₯ The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI TrendRadar
βAI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.π― εε«δΏ‘ζ―θΏθ½½οΌδ½ η
AI