mortogo321

FastAPI Claude Voice Agent — DevOps skill for Claude Code

DevOps community

FastAPI realtime voice AI agent with Claude Opus 4.7 (adaptive thinking + prompt caching + manual tool loop), Twilio Media Streams, Deepgram Nova-3 STT, ElevenLabs TTS, PostgreSQL, Docker, CI.

How to install FastAPI Claude Voice Agent

This entry records only its repository, not the path inside it, so there is no exact command to give. Open mortogo321/fastapi-claude-voice-agent and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What FastAPI Claude Voice Agent does

FastAPI realtime voice AI agent with Claude Opus 4.7 (adaptive thinking + prompt caching + manual tool loop), Twilio Media Streams, Deepgram Nova-3 STT, ElevenLabs TTS, PostgreSQL, Docker, CI.

Alternatives in DevOps

  • Appwrite — Appwrite® - complete cloud infrastructure for your web, mobile and AI apps 57.5k ★
  • Video Podcast Maker — AI-powered video podcast creation skill for coding agents 529 ★
  • Claude2api — Claude2API 是基于 Go + Docker 构建的 Claude.ai API 兼容网关、账号池与网页镜像服务,支持 OpenAI Chat Completions、Responses 和 Anthropic 71 ★

README

fastapi-claude-voice-agent

Production-ready realtime voice AI agent built on **FastAPI**, **Anthropic Claude Opus 5**, **Twilio Media Streams**, **Deepgram Nova-3 STT (speech-to-text)**, and **ElevenLabs TTS (text-to-speech)**. Handles inbound PSTN (public switched telephone network) calls and browser WebRTC (Web Real-Time Communication), runs an agentic tool-use loop with prompt caching and adaptive thinking, and persists every turn to PostgreSQL.

Features

  • PSTN voice in/out via Twilio Programmable Voice + Media Streams (μ-law 8kHz over WebSocket)
  • Browser WebRTC endpoint for low-latency demos
  • Streaming STT with Deepgram Nova-3 (multilingual: English + Thai)
  • LLM (large language model) Claude Opus 5 with adaptive thinking, prompt caching on the system prompt and tool definitions, and a manual agentic tool-use loop tuned for sub-second voice latency
  • Streaming TTS with ElevenLabs (eleven_turbo_v2_5)
  • Real barge-in — assistant playback runs as a cancellable asyncio task; a partial user transcript interrupts it mid-chunk
  • Tool use — appointment availability, booking, SMS (short message service) confirmation
  • Persistence — call sessions, transcript turns, tool calls (PostgreSQL + SQLAlchemy 2.0 async + Alembic)
  • Hardening — Twilio webhook HMAC (hash-based message authentication code) validation, per-process CallGate caps concurrent calls, fail-fast settings validation in production
  • Observability — structlog JSON logs, request-ID middleware with contextvars correlation across async boundaries, per-stage latency (STT endpointing, LLM TTFT (time-to-first-token), tool call, TTS TTFB (time-to-first-byte))
  • Containerized — multi-stage Dockerfile (non-root, healthcheck), docker-compose for local dev (Postgres + Redis + one-shot Alembic migration)
  • CI (continuous integration) — code quality (ruff format + lint + mypy --strict), tests with coverage, docker build — GitHub Actions
  • **Typed pipeli