ExoQuery Bench — AI skill for Claude Code
Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL.
How to install ExoQuery Bench
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open RyanSingh0/ExoQuery-Bench and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What ExoQuery Bench does
Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL. Scores Claude, OpenAI, Gemini and local models with RAG, self-correcting agents and AstroFetch's MCP tools on the NASA Exoplanet Archive, and shows where answers fail.
Alternatives in AI
- WeKnora — Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, an 26.9k ★
- CLI 代理 API — English 中文 一个为 CLI 提供 OpenAI/Gemini/Claude/Codex 兼容 API 接口的代理服务器 19k ★
- WindsurfAPI — Turn Windsurf / Devin Desktop's 100+ AI models (Claude, GPT, Gemini, DeepSeek, Kimi, GLM, SWE) into OpenAI-, A 3k ★
README
ExoQuery-Bench
**How accurately do LLMs answer plain-English questions from NASA astronomy archives?**
ExoQuery-Bench is an open benchmark and evaluation harness for natural-language-to-ADQL assistants on the [NASA Exoplanet Archive](https://exoplanetarchive.ipac.caltech.edu/). It scores Claude, OpenAI, Gemini and local models across retrieval (RAG), self-correcting agents and the public [AstroFetch](https://astrofetch.ipac.caltech.edu/) MCP tools, and reports *where* answers fail, not just how often.
**Status: work in progress (started October 2026).** No results are published yet. The roadmap below shows what exists and what is next. Numbers in this README describe the archive itself, checked live on 2026-10-07.
Why this exists
Assistants such as AstroFetch turn a question like *"How many planets has the transit method found?"* into ADQL, run it against the archive's TAP service and explain the result. The [AstroFetch FAQ](https://astrofetch.ipac.caltech.edu/faq) says the most common errors are wrong column names and unit mismatches, and that quality is currently judged through user ratings. A public, reproducible score would show which models, prompts and tools reduce those errors.
The archive has real traps for a language model. Each one below was checked live against the TAP service on 2026-10-07 and against the archive's [column definitions](https://exoplanetarchive.ipac.caltech.edu/docs/API_PS_columns.html) and [TAP guide](https://exoplanetarchive.ipac.caltech.edu/docs/TAP/usingTAP.html):
| Trap | What a model might write | What happens |
|---|---|---|
| One row per solution, not per planet | select count(*) from ps where discoverymethod='Transit' |
36,100 rows instead of 4,711 planets. The ps table holds 40,194 rows for 6,375 planets; only default_flag=1 (or the pscomppars table) gives one row per planet. |
| Earth vs Jupiter radii | pl_radj when the question means Earth radii |
Answers off by 11.2× (1 RJup</ |
Related Skills
Vibe Explain
Cognitive debt map for AI-generated code. Surfaces opaque blocks you don't fully own or understand, generates
People Search Bench
The first open benchmark for evaluating AI-powered people search agents
Mythos Bench
Jagged Frontier: LLM vulnerability detection benchmark harnesses (API + Claude Code agentic)
10x Bench Kit
Create your own benchmark for coding AI Agents
Bench Watch
Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart p
Shellbench
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based sc
Related Agents
AI ML
AI/ML 통합 전문가 + LLM API 최신 모델/SDK 코딩 가이드. RAG 시스템, 문서 분석, OpenAI/Anthropic/Gemini/Ollama 최신 API 보장. "AI integra
Timps AI Workflow Orchestrator
Turn a plain-English multi-step workflow into executable code for LangGraph, Temporal, or Claude-Flow — plus a
Agy Stylist
Rewrites prose for style through Gemini via the agy CLI, in any register (legal, technical, academic, commerci