Webpage Text Extractor — AI skill for Claude Code
Dump any folder's text files (HTML, CSS, JS.
How to install Webpage Text Extractor
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open wachin/webpage-text-extractor and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Webpage Text Extractor does
Dump any folder's text files (HTML, CSS, JS...) into one .txt — perfect for sending saved web pages to ChatGPT, Claude, or any AI agent.
Alternatives in AI
- Artifacts Builder — Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech 97.5k ★
- Repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k ★
- Context Dump — Dump current context for model switch or context limit recovery 508 ★
README
**🇪🇸 Spanish version available:** [README_ES.md](README_ES.md) — Versión en español para público de habla hispana.
webpage-text-extractor
**Extract all text from a folder — like a web page saved with `Ctrl + S` in Chrome — and concatenate it into a single `.txt` file ready to send to an AI agent.**
[](https://www.python.org/) [](LICENSE) [](#requirements)
Why this script exists?
When you ask an AI agent (ChatGPT, Claude, a local assistant, etc.) to analyze a web page, the most practical approach is often not to pass the URL but **the actual content of the page**: its HTML, styles, and structure.
The workflow is:
- Open the page in Google Chrome.
- Press
Ctrl + S(on macOS⌘ + S) and choose "Webpage, Complete". - Chrome creates a folder with the page's HTML and all its resources (CSS, JS, images).
- Run this script on that folder.
- Get a single
.txtfile with the readable content of all text files, clearly separated by headers.
That `.txt` file is much easier to handle: you attach it to your AI agent, paste it in a chat, or process it as you wish, without having to send dozens of loose files.
Features
- Recursive: traverses all subdirectories with
os.walk. - Text detection: attempts to read each file as UTF-8; binary files (images, fonts, videos) are listed by name without cluttering output.
- Clear separators: each text file is preceded by a header with its full path:
--- Text file content: Blog-example/page.html --- - Zero dependencies: only Python standard library. No
pip installneeded. - Two versions:
extract_text.py— version with command-line arguments (recommended).extract_text_simple.py— original minimal version, edit two variables and run.
Requiremen
Related Skills
Video To LLM Context Extractor
Turn any video into a detailed, LLM-ready PDF document. Perfect for feeding visual and transcribed context int
Agent Ready SEO
A Claude Code skill for making a site readable and citable by AI answer engines as well as search crawlers: th
Hyperframes Cn
HyperFrames 是一个开源框架,可将 HTML、CSS、媒体与可定位(seekable)动画转化为确定性的 MP4 视频。你可以在本地通过 CLI 使用它,让 AI 编程智能体借助 skills 使用它,或者把它
Html2elementor
Convert HTML + CSS to Elementor JSON (importable in WordPress). Works as a Claude Code / openclaw skill or sta
Rogue Web Artifacts Builder Skill
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
MCP Memory Sqlite
A personal knowledge graph and memory system for AI assistants using SQLite with optimized text search. Perfec
Related Agents
Canvas To Code Extractor
Decodes any design-source HTML (Claude Design bundle, Figma Export-to-Code, V0, Lovable, Webflow, generic HTML
MVP Builder
Full-stack MVP builder. Takes a plain-English description and delivers an award-winning static site — no frame
UI Design Reviewer
Use PROACTIVELY to review generated HTML/CSS mock files against a DESIGN.md specification. MUST BE USED after