PDF Audio Video To Markdown With AI — AI skill for Claude Code
Convert PDF, Audio, Video to Markdown with smart processing.
How to install PDF Audio Video To Markdown With AI
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open evan966890/PDF-Audio-Video-to-Markdown-with-AI and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What PDF Audio Video To Markdown With AI does
Convert PDF, Audio, Video to Markdown with smart processing.
Alternatives in AI
- Geo SEO Claude — GEO-first SEO skill for Claude Code 3.6k ★
- Generative Media Skills — Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI) 3k ★
- Qwen Audio Agent — A realtime voice runtime that keeps Agents talking, working, and present 2.3k ★
README
PDF-Audio-Video-to-Markdown with AI
**Universal AI Skill** for Claude Code / Cursor / Antigravity / Windsurf and more
Let AI handle everything - from setup to transcription!
**通用 AI 技能** 适用于 Claude Code / Cursor / Antigravity / Windsurf 等
让 AI 处理一切 - 从环境配置到文件转录!
[](https://claude.ai) [](https://cursor.com) [](https://antigravity.dev) [](https://windsurf.ai)
[](https://www.python.org/) [](https://opensource.org/licenses/MIT) [](https://github.com/evan966890/PDF-Audio-Video-to-Markdown-with-AI)
**English** | **中文**
Intelligently convert PDF, Audio, Video & Images to Markdown text, especially optimized for **meeting recordings transcription**.
将 PDF / 音频 / 视频 / 图像 **智能转换** 为 Markdown 文本,特别适合 **会议录屏转文字** 场景。
Table of Contents / 目录
- Why This Tool / 为什么选择
- Features / 功能特性
- Privacy & Security / 隐私安全
- System Requirements / 系统要求
- Quick Start / 快速开始
- Usage / 使用方法
- FAQ / 常见问题
- Roadmap / 路线图
- Contributing / 贡献
Why This Tool / 为什么选择
| Pain Point 痛点 | Solution 解决方案 |
|---|---|
| Meeting recordings unsearchable / 会议录屏无法搜索 | Auto-transcribe to searchable Markdown / 自动转写为可搜索 Markdown |
| Scanned PDF tex |
Related Skills
Whisperx Transcribe
Claude Code skill/plugin: transcribe audio & video into clean, speaker-labeled Markdown via WhisperX, so an LL
Nutrient Agent Skill
Document processing with Nutrient DWS API: convert (PDF/DOCX/XLSX/PPTX/HTML/images), extract text/tables, OCR
Multimodal Voice Assistant
This project is a multi-modal AI voice assistant that uses LM Studio, OpenAI API or Claude Code, audio process
Ezpzfile MCP
File tools for AI agents. Read and convert documents (DOCX, PDF, HWP, HWPX, EML), edit PDFs, resize, clean and
Review Video
Review a rendered video with AI vision. Uses TwelveLabs to watch the output, analyze pacing/composition/audio
Claude Shorts
Interactive longform-to-shortform video creator — Claude Code skill with Remotion-rendered animated captions,
Related Agents
Voice Engineer
GAIA voice interaction specialist. Use PROACTIVELY for Whisper ASR, Kokoro TTS, the Talk SDK, speech-to-speech
Dsp Agent
Implement audio processing and DSP algorithms for Stage 2. Use PROACTIVELY after foundation-shell-agent comple
Frame Describer
Describes video frames as detailed text. Used when frame_mode is "descriptions" to convert visual frames into