Video Talkcraft banner
Vincentwei1021 Vincentwei1021

Video Talkcraft

Design community

Description

Agent skill that turns Claude Code / Codex into a motion-design studio for narration videos — word-level voiceover sync, 78 motion recipe cards, an anti-slideshow camera system, Remotion rendering.

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

video-talkcraft

[![Gallery](https://img.shields.io/badge/Gallery-live%20previews-7A5AF8)](https://vincentwei1021.github.io/video-talkcraft/) [![License](https://img.shields.io/badge/License-PolyForm%20Noncommercial-blue)](LICENSE)

**An agent skill for crafting high-quality narration videos: word-level voiceover sync · 78 motion recipe cards · a 7-layer anti-slideshow shot system · triple-gate QA**

[English](README.md) | [中文](README_CN.md)

**video-talkcraft** is the narration-video installment of the [video-shotcraft](https://github.com/Vincentwei1021/video-shotcraft) series: an AI agent skill that turns Claude Code or Codex into a motion-design studio for talking videos. Hand it a narration script and a finished voiceover, and it aligns word-level timestamps locally, storyboards every semantic beat into a SHOTBOOK, then renders a high-quality explainer with [Remotion](https://www.remotion.dev/) — kinetic type, evidence screenshots, camera moves, plain-cut subtitles, and film-grade SFX, all locked to the voice.

The methodology docs and recipe cards are written in Chinese — the toolkit is built Chinese-narration-first (mixed Chinese/English narration is fully supported). Agents read them natively.

🖼️ [**Browse all 78 motion previews in the live Gallery »**](https://vincentwei1021.github.io/video-talkcraft/)

✨ Highlights

  • Word-level voiceover syncscripts/timestamps_cpu.py aligns your script to the audio (FireRedASR2-CTC int8 by default, faster-whisper as the zero-download fallback). Benchmarked against a GPU forced aligner on a 110s mixed-language narration: median per-character offset 20–40 ms, worst case 200 ms, zero false QA flags. Every motion beat anchors to the exact word.
  • 78 motion recipe cards — each with intent, parameters, known pitfalls, and a runnable HTML preview — browse them all in the online Gallery or locally wit