techdou

Boson AI Skill — AI skill for Claude Code

AI community

ZCode/Claude Code agent skill for Boson AI — Higgs TTS 3 (102-language speech synthesis) and Higgs Avatar (talking-head video).

How to install Boson AI Skill

This entry records only its repository, not the path inside it, so there is no exact command to give. Open techdou/boson-ai-skill and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Boson AI Skill does

ZCode/Claude Code agent skill for Boson AI — Higgs TTS 3 (102-language speech synthesis) and Higgs Avatar (talking-head video).

Alternatives in AI

  • EchoBird — One-click install + model switch:Claude Code,Codex CLI (OpenAI), Grok Build (xAI), DeepSeek Harness, Kimi Code 3.1k ★
  • Qwen Audio Agent — A realtime voice runtime that keeps Agents talking, working, and present 2.3k ★
  • AI Avatar System — 🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip 422 ★

README

Boson AI Skill | Boson AI 语音与数字人 Skill

[English](#english) | [中文](#中文)


中文

可移植 Agent Skill,调用 [Boson AI](https://boson.ai) 的 **Higgs TTS 3** 和 **Higgs Avatar** 模型。支持 102 种语言的语音合成、声音克隆、情感/风格控制、数字人口播视频生成,全部通过命令行完成。

触发场景

  • 文本转语音(TTS)、多语言配音
  • 声音克隆、参考音频注册复用
  • 数字人/口播视频、音频转视频
  • 情感标签控制(内联 emotion tag)

下方英文为完整文档。


English

A [ZCode](https://zcode.ai) / Claude Code / OpenCode agent skill for [Boson AI](https://boson.ai)'s **Higgs TTS 3** and **Higgs Avatar** models. Generate speech in 102 languages, clone voices, create talking-head avatar videos — all from the command line.

Features

**Higgs TTS 3 — Text-to-Speech**

  • 102 languages with single-digit WER/CER (Chinese, English, Japanese, Korean, Thai, Vietnamese, French, German, Spanish, Arabic, Hindi, and more)
  • 6 preset voices + instant voice cloning from reference audio
  • Inline emotion / style / sound-effect / prosody tags for fine-grained control
  • Streaming PCM for low-latency output
  • Batch synthesis from segments JSON

**Higgs Avatar — Talking-Head Video**

  • Animate a still photo with driving audio (audio-to-video)
  • Or synthesize voice + animate in one request (text-to-video)
  • Streaming fMP4 for real-time playback
  • Auto-fallback when text-to-video hits the known server-side bug

**Self-Update Guard**

  • Runtime check against Boson's official openapi.json and .md docs
  • Detects model deprecation, voice list changes, and parameter drift
  • 24h cache, non-blocking by default

Quick Start

Prerequisites

  • Python 3.9+
  • requests library: pip install requests
  • A Boson AI API key (bai-...) from boson.ai

Install

Clone this repo into your agent skills directory:

# For ZCode / Claude Code / OpenCode
git clone https://github.com/techdou/boson-ai-skill.git ~/.agents/skills/boson-ai-skill

Set API Key

export BOSON_API_KEY="bai-your-key-here"

Windows PowerShell:

$env:BOSON_A