Awesome Scientific LLM Benchmarks banner
subinium subinium

Awesome Scientific LLM Benchmarks

AI community

Description

A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

Awesome Scientific LLM Benchmarks [![Awesome](https://awesome.re/badge.svg)](https://awesome.re) [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

Benchmarks for evaluating large language models on **scientific reasoning and discovery** — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.

_Listed alphabetically within each domain._

Contents

General / Multi-domain Science

Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.

Benchmark Org Year Paper Code Stars Description
AGIEval Microsoft 2023 paper code ![stars](https://img.shields.io/github/stars/ruixiangcui/AGIEval?style=flat&logo=github&logoColor=white&label=&color=181717) Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.
ARB DuckAI / Georgia Tech 2023 paper code ![stars](https://img.shields.io/github/stars/TheDuckAI/arb?style=flat&logo=github&logoColor=white&label=&color=181717) Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.
ARC (AI2 Reasoning Challenge) Allen AI (AI2) 2018 paper [code](http