AIBit-Discover Open Source Projects AIBit-Discover Open Source Projects
Open Source ProjectsWeb Scraping & DataAI Agents & AutomationAI Tools & Resources
More
Learning & TutorialsAI Research & BenchmarksDevelopment & SecurityWeb & InfrastructureMedia & Content CreationHardware & Edge AIStartup Resources
AIBit-Discover Open Source Projects › AI Tools & Resources› Voice & Audio AI

September 19, 2026

VoiceStudio: A Fully Local, Open-Source Voice AI Workbench

Explore VoiceStudio, a local ElevenLabs alternative for voice cloning, dubbing, dictation, transcription, and audiobooks across 646 languages.

  • Sep 14, 2026

    YuE2: Open Music Generation with Editable Scores and AI Agents

    YuE2 turns lyrics and style prompts into editable musical scores, frontier-quality songs, zero-shot covers, and agent-driven revisions.

  • Sep 14, 2026

    VoiceStudio: A Fully Local, Open-Source ElevenLabs Alternative

    VoiceStudio brings voice cloning, dubbing, dictation, transcription, and audiobooks to your hardware with 16 TTS and 11 ASR engines.

  • Aug 30, 2026

    StemDeck: Free, Local 6-Stem AI Audio Separation with Demucs and Tauri

    StemDeck is an open-source, local-first stem extraction studio that isolates vocals, drums, bass, guitar, and piano with no cloud subscriptions or uploads.

  • Aug 14, 2026

    HOT-Step CPP: Local AI Music Generation with C++ and GGML

    Generate stereo 48 kHz music locally with HOT-Step CPP, a feature-rich C++/GGML interface for ACE-Step, audio tools, plugins, and model training.

  • Aug 2, 2026

    Inflect v2: Local Text-to-Waveform TTS at 3.96M and 9.36M Parameters

    Inflect v2 delivers complete 24 kHz English TTS in two compact models (3.96M and 9.36M params) with no external vocoder. Explore architecture, evaluation, and CPU deployment.

  • Aug 2, 2026

    KittenTTS: State-of-the-Art TTS Under 25MB That Runs on CPU

    KittenTTS delivers high-quality text-to-speech in models as small as 25MB, running entirely on CPU. Learn how to use it, its API, and why it matters for edge AI.

  • Jul 30, 2026

    CrispASR: One C++ Binary for 53 ASR and 51 TTS Models

    CrispASR is a unified C++ speech engine built on ggml, supporting 53 ASR and 51 TTS models with zero Python dependencies. Run Cohere Transcribe, Parakeet, Voxtral, Qwen3, and more from a single CLI.

  • Jul 22, 2026

    FluidAudio: Run SOTA Audio AI Models Locally on Apple Devices with CoreML

    FluidAudio is a Swift SDK for running state-of-the-art speech-to-text, text-to-speech, speaker diarization, and voice activity detection entirely on-device using Apple's Neural Engine. This post explores its capabilities, architecture, and how to integrate it into your apps.

  • Jul 22, 2026

    SpeechRecognition: A Comprehensive Python Library for Speech-to-Text

    Explore SpeechRecognition, a versatile Python library supporting multiple speech recognition engines and APIs, both online and offline, with practical examples and troubleshooting tips.

  • Jul 11, 2026

    Speech Swift: On-Device AI Speech Toolkit for Apple Silicon

    Explore Speech Swift, an open-source toolkit for ASR, TTS, speech-to-speech, VAD, and diarization on Apple Silicon using MLX and CoreML.

  • Jun 6, 2026

    Miso TTS 8B: A High-Quality Open-Source Text-to-Speech Model

    Miso TTS 8B is a state-of-the-art, open-source text-to-speech model with 8 billion parameters, offering highly emotive speech generation and voice cloning capabilities.

Previous 1 / 4 Next

Curated AI tools, open source projects, tutorials, and resources for developers building with artificial intelligence.

Terms of Service Privacy Policy © 2026 AIBit-Discover Open Source Projects