wachin

Webpage Text Extractor — AI skill for Claude Code

AI community

Dump any folder's text files (HTML, CSS, JS.

How to install Webpage Text Extractor

This entry records only its repository, not the path inside it, so there is no exact command to give. Open wachin/webpage-text-extractor and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Webpage Text Extractor does

Dump any folder's text files (HTML, CSS, JS...) into one .txt — perfect for sending saved web pages to ChatGPT, Claude, or any AI agent.

Alternatives in AI

  • Artifacts Builder — Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech 97.5k ★
  • Repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file 22.7k ★
  • Context Dump — Dump current context for model switch or context limit recovery 508 ★

README

**🇪🇸 Spanish version available:** [README_ES.md](README_ES.md) — Versión en español para público de habla hispana.

webpage-text-extractor

**Extract all text from a folder — like a web page saved with `Ctrl + S` in Chrome — and concatenate it into a single `.txt` file ready to send to an AI agent.**

[![Python](https://img.shields.io/badge/Python-3.6%2B-blue.svg)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE) [![No dependencies](https://img.shields.io/badge/dependencies-none-orange.svg)](#requirements)


Why this script exists?

When you ask an AI agent (ChatGPT, Claude, a local assistant, etc.) to analyze a web page, the most practical approach is often not to pass the URL but **the actual content of the page**: its HTML, styles, and structure.

The workflow is:

  1. Open the page in Google Chrome.
  2. Press Ctrl + S (on macOS ⌘ + S) and choose "Webpage, Complete".
  3. Chrome creates a folder with the page's HTML and all its resources (CSS, JS, images).
  4. Run this script on that folder.
  5. Get a single .txt file with the readable content of all text files, clearly separated by headers.

That `.txt` file is much easier to handle: you attach it to your AI agent, paste it in a chat, or process it as you wish, without having to send dozens of loose files.

Features

  • Recursive: traverses all subdirectories with os.walk.
  • Text detection: attempts to read each file as UTF-8; binary files (images, fonts, videos) are listed by name without cluttering output.
  • Clear separators: each text file is preceded by a header with its full path:
    --- Text file content: Blog-example/page.html ---
  • Zero dependencies: only Python standard library. No pip install needed.
  • Two versions:
    • extract_text.py — version with command-line arguments (recommended).
    • extract_text_simple.py — original minimal version, edit two variables and run.

Requiremen