cyc05

Video OCR Storyboard Skill — Design skill for Claude Code

Design community

A Trae/Claude skill for extracting on-screen text from videos without ffmpeg.

How to install Video OCR Storyboard Skill

This entry records only its repository, not the path inside it, so there is no exact command to give. Open cyc05/video-ocr-storyboard-skill and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Video OCR Storyboard Skill does

A Trae/Claude skill for extracting on-screen text from videos without ffmpeg. Per-second frame sampling with OpenCV, batch OCR via RapidOCR (Chinese + English), labeled contact sheets for visual verification of ambiguous data, and a workflow for building second-accurate video timelines and voice-over scripts.

Alternatives in Design

  • Canvas — Open, create, or update a visual canvas — add images, text, PDFs, wiki pages, and banana-generated assets to O 3.3k ★
  • Stitch Remotion — Generate walkthrough videos from app designs 2.6k ★
  • Clipify — Claude Code skill: turn long videos into social-ready clips 549 ★

README

video-ocr-storyboard-skill

A Trae/Claude skill for extracting on-screen text from videos — **no ffmpeg required**. Per-second frame sampling with OpenCV, batch OCR via RapidOCR (Chinese + English), labeled contact sheets for visual verification of ambiguous data, and a workflow for building second-accurate video timelines and voice-over scripts.

Originally battle-tested on a 3-minute, 180-frame data-driven documentary ("From 0 to...", a China-through-data timeline of 50+ tech milestones), where every on-screen figure had to be extracted and verified second by second.

What it does

Step Script Output
1. Probe & sample scripts/extract_frames.py One JPG per second (f_0000.jpg, ...), plus fps / resolution / duration
2. Batch OCR scripts/batch_ocr.py JSON timeline: second → recognized text lines (top-to-bottom, left-to-right)
3. Visual verification scripts/make_contact_sheets.py Labeled 2×N montage sheets to double-check huge numbers & cross-dissolve frames
4. Timeline & script guided by SKILL.md Card-level timeline with entry/exit timecodes, ready for narration writing

Why not just "run OCR on everything"?

  • Cross-dissolves lie: transition frames blend two adjacent data cards — ghost text belongs to the previous or next card. Only trust a number on the frame where it is fully formed.
  • Huge semi-transparent digits are OCR-unfriendly; the contact-sheet pass catches what raw OCR mangles.
  • Constant noise: channel logos and watermarks appear on every frame — the workflow treats them as background, not content.
  • Never silently "fix" data: if on-screen numbers contradict known facts, transcribe what the video shows and flag the discrepancy.

Requirements

pip install opencv-python pillow rapidocr-onnxruntime

First RapidOCR run downloads ONNX models automatically (needs network once).

Usage

# 1. Extract one frame per second (optionally a segm