RyanSingh0

ExoQuery Bench — AI skill for Claude Code

AI community

Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL.

How to install ExoQuery Bench

This entry records only its repository, not the path inside it, so there is no exact command to give. Open RyanSingh0/ExoQuery-Bench and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What ExoQuery Bench does

Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL. Scores Claude, OpenAI, Gemini and local models with RAG, self-correcting agents and AstroFetch's MCP tools on the NASA Exoplanet Archive, and shows where answers fail.

Alternatives in AI

  • WeKnora — Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, an 26.9k ★
  • CLI 代理 API — English 中文 一个为 CLI 提供 OpenAI/Gemini/Claude/Codex 兼容 API 接口的代理服务器 19k ★
  • WindsurfAPI — Turn Windsurf / Devin Desktop's 100+ AI models (Claude, GPT, Gemini, DeepSeek, Kimi, GLM, SWE) into OpenAI-, A 3k ★

README

ExoQuery-Bench

**How accurately do LLMs answer plain-English questions from NASA astronomy archives?**

ExoQuery-Bench is an open benchmark and evaluation harness for natural-language-to-ADQL assistants on the [NASA Exoplanet Archive](https://exoplanetarchive.ipac.caltech.edu/). It scores Claude, OpenAI, Gemini and local models across retrieval (RAG), self-correcting agents and the public [AstroFetch](https://astrofetch.ipac.caltech.edu/) MCP tools, and reports *where* answers fail, not just how often.

**Status: work in progress (started October 2026).** No results are published yet. The roadmap below shows what exists and what is next. Numbers in this README describe the archive itself, checked live on 2026-10-07.


Why this exists

Assistants such as AstroFetch turn a question like *"How many planets has the transit method found?"* into ADQL, run it against the archive's TAP service and explain the result. The [AstroFetch FAQ](https://astrofetch.ipac.caltech.edu/faq) says the most common errors are wrong column names and unit mismatches, and that quality is currently judged through user ratings. A public, reproducible score would show which models, prompts and tools reduce those errors.

The archive has real traps for a language model. Each one below was checked live against the TAP service on 2026-10-07 and against the archive's [column definitions](https://exoplanetarchive.ipac.caltech.edu/docs/API_PS_columns.html) and [TAP guide](https://exoplanetarchive.ipac.caltech.edu/docs/TAP/usingTAP.html):

Trap What a model might write What happens
One row per solution, not per planet select count(*) from ps where discoverymethod='Transit' 36,100 rows instead of 4,711 planets. The ps table holds 40,194 rows for 6,375 planets; only default_flag=1 (or the pscomppars table) gives one row per planet.
Earth vs Jupiter radii pl_radj when the question means Earth radii Answers off by 11.2× (1 RJup</