Awesome Scientific LLM Benchmarks
Description
A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.
Installation
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open the source below and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
README
Awesome Scientific LLM Benchmarks [](https://awesome.re) [](LICENSE)
Benchmarks for evaluating large language models on **scientific reasoning and discovery** — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.
_Listed alphabetically within each domain._
Contents
- General / Multi-domain Science
- Mathematics
- Physics & Astronomy
- Chemistry
- Materials Science
- Biology & Life Sciences
- Agentic Science & AI Research
General / Multi-domain Science
Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| AGIEval | Microsoft | 2023 | paper | code |  | Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions. |
| ARB | DuckAI / Georgia Tech | 2023 | paper | code |  | Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric. |
| ARC (AI2 Reasoning Challenge) | Allen AI (AI2) | 2018 | paper | [code](http |
Related Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Awesome Llm Apps
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI