Free, open-source curriculum for making money with generative AI image, video, and audio — for creators and agencies.
-
Updated
Aug 2, 2026
Free, open-source curriculum for making money with generative AI image, video, and audio — for creators and agencies.
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
Full-featured DAW, DJ, and VJ app interoperable w/Ableton, Reaper, Resolume & more. Stable Audio 3, Magenta RT2, Suno API, Chimera track fusion, Demucs stems, MIDI generate/notate, img > spectrogram > music, drawing > music, VST3 & .gan plugins, automix & key-lock, GLSL shaders, volumetric video, Quest 3 XR interface, MIDI auto-map, RAG assistant
Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.
Draft to Take beta: local-first AI audio production studio powered by IndexTTS2, Docker, Qwen, OmniVoice, SFX, ambience, and music sidecars.
Free real-time AI Noise Gate VST3/AU plugin. Removes coughs, sneezes, and other artifacts from your live streams, podcasts, and videos.
ARCHIVED USE theDAW at https://github.com/gantasmo/theDAW - Browser-based AI audio DAW for Stable Audio 3 with text-to-audio, inpainting, LoRA training, FFmpeg effects, waveform editing, sequencer, piano roll, and persistent library.
Real-Time Deepfake Pipeline
AI-powered audio manipulation studio — sample pack creation, algorithmic composition, text-to-audio generation, and ChatGPT on one screen
Local-first CLI that turns Markdown scripts into multi-speaker podcast-style audio using Coqui XTTS v2.
Music Generation Using Deep Learning🎶🎵
Community list of AI tools for audio and music
AI Audio Content Creation Platform for Podcasts, Narration, Voice Generation and Audio Production. Create Professional Audio Content from Text with Modern Audio Workflows.
AI Voice Agents: Exploring the Next Generation of Human-Machine Interaction! 🎙️🤖🎧
A local-first EPUB reader with high-fidelity neural text-to-speech, word-level synchronization, and Next.js/FastAPI/ONNX stack.
AudioInsight is a web application that processes audio, generates transcriptions, and allows users to ask questions about the related audio.
High-performance KittenTTS API server with a built-in web UI, OpenAI-compatible routes, long-form text support, and optional CUDA acceleration.
An approach to Andrej Karpathy's LLM challenge, as outlined here: https://twitter.com/karpathy/status/1760740503614836917
Interactive bilingual audio technology sharing site with waveform labs, DSP, hardware, software, AI audio, and application scenarios.
Add a description, image, and links to the ai-audio topic page so that developers can more easily learn about it.
To associate your repository with the ai-audio topic, visit your repo's landing page and select "manage topics."