Audio middleware
-
AI-Audio-Content-Creator
Transform text into professional audio and podcasts using automated voice synthesis tools.
-
ai-media-studio-cli
🎨 Generate stunning videos, images, and music with ease using Google's AI models and simple text prompts in the AI Media Studio CLI.
-
docker-talkies
OpenAI-compatible audio server in Docker. 7 ASR backends (Whisper, Distil-Whisper, Parakeet, Canary, Canary-Qwen) + 2 TTS engines (Kokoro, Qwen3-TTS voice cloning). Single /v1/audio/{transcriptions,speech,voices} surface. CPU + CUDA images. Hot model swap. MCP server built in.
-
Lattice
CLI toolkit for music collectors: library trees, integrity checks, cover art, tag audits.
-
LivePilot
Agentic production system for Ableton Live 12 — 467 tools across 56 domains. Device atlas, user corpus, Splice intelligence, 9-band spectral perception, Creative Director, semantic moves, and 12 creative engines.
-
parakeet-tdt-0.6b-v2-Batch-Transcriber
🎙️ Transcribe audio efficiently with Parakeet Batch Transcriber, producing accurate, timestamped SRT files from long recordings like podcasts and lectures.
-
qobuz-librarian
Build and maintain a complete, lossless library from Qobuz: gap-fill, quality upgrades, repair, and clean beets imports. Web UI or CLI, one Docker image.
Audio Python -
QQgroup-annual-report-analyzer
📊 Analyze QQ group chat records to create stunning annual reports with customizable features and AI integration for enhanced insights.
-
stripped-vocal-remover-windows
The easiest way to install the Stripped Vocal Remover on Windows. This auto-installer fixes dependency errors, bypasses Hugging Face logins, and provides a beginner-friendly local UI for audio separation.
-
TDPilot
TDPilot v2.1.0 — TouchDesigner AI assistant (114 MCP tools, correctness-first brain: plan -> execute -> validate -> rollback, 656 operator cards, read-only cockpit UI)
-
TrackSplit
Turn chaptered videos into gapless, tagged albums for Jellyfin, Lyrion, Navidrome, and Plex
-
VoiceGuard
🔒 Enhance security with VoiceGuard, an AI-driven voice authentication system powered by OpenAI’s ChatGPT and Whisper for reliable voice identification.