Phantom Agent — AI skill for Claude Code
Autonomous AI agent for BitGN PAC1 Challenge — ~86% score with 12 hot-reloadable skills, live dashboard, and self-correcting classification.
How to install Phantom Agent
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open vakovalskii/phantom-agent and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Phantom Agent does
Autonomous AI agent for BitGN PAC1 Challenge — ~86% score with 12 hot-reloadable skills, live dashboard, and self-correcting classification.
Alternatives in AI
- WeKnora — Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, an 26.9k ★
- Pro Workflow — 1,400+ Battle-tested Claude Code workflows from power users 1.4k ★
- Phantom — An AI co-worker with its own computer 1.3k ★
README
Phantom — Autonomous Agent for BitGN PAC1 Challenge
[Русский](docs/README_RU.md) | [中文](docs/README_ZH.md)
An autonomous file-system agent built with [OpenAI Agents SDK](https://github.com/openai/openai-agents-python) that solves the [BitGN PAC1 Challenge](https://bitgn.com/challenge/PAC) — a benchmark for AI agents operating in sandboxed virtual environments.
**Current score: ~86% (37/43 tasks)**


What is PAC1?
[BitGN](https://bitgn.com) runs agent benchmarks where autonomous agents solve real-world tasks inside isolated sandbox VMs. Each task gives the agent a file-system workspace and a natural language instruction. The agent must explore, reason, and execute — no human in the loop.

PAC1 covers 43 tasks across:
- CRM operations — lookups, email sending, invoice handling
- Knowledge management — capture, distill, cleanup
- Inbox processing — with prompt injection traps and OTP verification
- Security — detecting and denying hostile payloads
Learn more: [bitgn.com/challenge/PAC](https://bitgn.com/challenge/PAC)
Architecture
User task → LLM Classifier (picks skill) → Agent(system_prompt + skill_prompt + task)
→ ReAct loop: LLM → tool call → result → LLM → ... → report_completion
- 12 specialized skills with hot-reloadable prompts (edit
.mdfiles, no restart needed) - Dual classifier — LLM-first with regex fallback and override logic
- Self-correcting agent — can call
list_skills/get_skill_instructionsto switch workflows mid-task - Auto grounding refs — tracks read/written files, injects references if model forgets
- Retry on empty — retries up to 3x if model returns text without tool calls
- Live dashboard — React + Vite with SSE streaming, heatmap compare, token tracking
Quick Start
Prerequisites
- Python 3.12+
- [uv](
Related Skills
Agent Arena
Open-source evaluation framework for LLM agents. Run head-to-head A/B tests, score with LLM-as-judge rubrics,
Lumbergh
Self-hosted web dashboard for supervising multiple Claude Code AI sessions in tmux. Live terminals, real-time
Verdandi
VERÐANDI — The Norn of Becoming. AI nervous system for real-time cross-instance awareness via Unix domain sock
ExoQuery Bench
Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL. Scores Claude, OpenAI
Codecrafters Claude Code Cpp
A high-performance C++ implementation of an autonomous AI coding agent inspired by Claude Code, built as part
ProductsFactory
Autonomous AI dev factory: sandboxed Claude Code agents design, build, review, and ship features across repos
Related Agents
Health Checker
Code health dashboard agent - structured 3-phase methodology (detect, measure, prescribe) with weighted 0-10 s
Role Scout
Find live job openings that match the user's target roles, score each against the parsed CV, and write ranked
Claude Code Hook Comms (hcom)
by aannoo - Lightweight CLI tool for real-time communication between Claude Code sub agents using hooks. Enabl