Multi Modal Agent TS — DevOps skill for Claude Code
TypeScript multimodal AI agent: GPT-4o / Claude / Gemini + Whisper + Ollama (LLaVA).
How to install Multi Modal Agent TS
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open laoposkj/multi-modal-agent-ts and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Multi Modal Agent TS does
TypeScript multimodal AI agent: GPT-4o / Claude / Gemini + Whisper + Ollama (LLaVA). REST API, streaming, Docker. Vision, audio & text in one flow.
Alternatives in DevOps
- Claudable — Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code, Codex, Gemini CL 4k ★
- Azure Functions TypeScript — Azure Functions for TypeScript 1.8k ★
- Codeg — Aggregate and browse AI coding agent sessions (Claude Code, Codex, Gemini CLI, etc.) in one place 850 ★
README
multi-modal-agent-ts
**Vision + text + audio TypeScript AI agent** Clone → `npm install` → `npm run dev` (no build step) → process any input modality. Use `npm run build` before `npm start` or when importing from `./dist`.
[](https://www.typescriptlang.org/) [](https://sdk.vercel.ai/) [](https://opensource.org/licenses/MIT) [](https://nodejs.org/)
What is this?
`multi-modal-agent-ts` is a **TypeScript agent that can combine images, audio, and text** in one flow. It uses multimodal chat models (OpenAI GPT‑4o, Anthropic Claude 3.x, Google Gemini) via the [Vercel AI SDK](https://sdk.vercel.ai/), OpenAI **Whisper** for speech-to-text (`whisper-api`), optional **local Whisper** via `@xenova/transformers` (`whisper-local`), and **Ollama** (e.g. LLaVA) for local vision without an API key (`local/llava`).
You can:
- Analyze images — paths,
httpsURLs, or buffers; optionalimageDetailfor OpenAI. - Transcribe audio — file path, URL, or buffer; Whisper API or optional local ASR.
- Fuse modalities — one user message with text + transcripts + images.
- Stream answers —
processStream()with transcript / vision / answer chunks. - Run an HTTP API —
POST /processwith JSON or multipart uploads.
Installation
cd multi-modal-agent-ts
npm install
npm run dev— runs the API withtsx(TypeScript directly; nodist/needed).npm run build— needed fornpm start, Docker, or importing from./dist/....
Requires **Node.js 20+**. For video frame/audio extraction and microphone capture helpers, install **ffmpeg** and use a normal audio device (see [Audio / mic](#audio--mic)).
Copy `.env.example` to `.env` and set keys for the provid
Related Skills
LibreChat
Enhanced ChatGPT Clone: Features OpenAI, Assistants API, Azure, Groq, GPT-4 Vision, Mistral, Bing, Anthropic,
Claude Proxy
这是一个部署在 Cloudflare Workers 上的 TypeScript 项目,它充当一个代理服务器,能够将 Claude API 格式的请求无缝转换为 OpenAI API 格式。这使得任何与 Claude A
GPT Telegram Worker
A multi-model AI Telegram bot powered by Cloudflare Workers, supporting various APIs including OpenAI, Claude,
Nodejs Starter Template
Production-grade REST API starter — TypeScript, Express, PostgreSQL + Prisma, JWT auth, Zod validation, Vitest
Sandy
Sandboxed TypeScript runtime for AI coding agents to query AWS — full SDK access with in-sandbox aggregation,
Chat
chat web app for teams, sass with user management and ratelimit, support chatgpt(openai & azure), claude, gemi
Related Agents
Multimodal Engineer Md
Multi-modal AI specialist for AutoBot's advanced AI capabilities. Use for computer vision, voice processing, s
Delegate
Expert LLM delegation specialist that seamlessly connects to external language models including GPT-4, GPT-3.5
Voice Engineer
GAIA voice interaction specialist. Use PROACTIVELY for Whisper ASR, Kokoro TTS, the Talk SDK, speech-to-speech