AIBit-Discover Open Source Projects AIBit-Discover Open Source Projects
Open Source ProjectsWeb Scraping & DataAI Agents & AutomationAI Tools & Resources
More
Learning & TutorialsAI Research & BenchmarksDevelopment & SecurityWeb & InfrastructureMedia & Content CreationHardware & Edge AIStartup Resources
AIBit-Discover Open Source Projects › AI Tools & Resources› Voice & Audio AI

August 14, 2026

HOT-Step CPP: Local AI Music Generation with C++ and GGML

Generate stereo 48 kHz music locally with HOT-Step CPP, a feature-rich C++/GGML interface for ACE-Step, audio tools, plugins, and model training.

  • Aug 2, 2026

    Inflect v2: Local Text-to-Waveform TTS at 3.96M and 9.36M Parameters

    Inflect v2 delivers complete 24 kHz English TTS in two compact models (3.96M and 9.36M params) with no external vocoder. Explore architecture, evaluation, and CPU deployment.

  • Aug 2, 2026

    KittenTTS: State-of-the-Art TTS Under 25MB That Runs on CPU

    KittenTTS delivers high-quality text-to-speech in models as small as 25MB, running entirely on CPU. Learn how to use it, its API, and why it matters for edge AI.

  • Jul 30, 2026

    CrispASR: One C++ Binary for 53 ASR and 51 TTS Models

    CrispASR is a unified C++ speech engine built on ggml, supporting 53 ASR and 51 TTS models with zero Python dependencies. Run Cohere Transcribe, Parakeet, Voxtral, Qwen3, and more from a single CLI.

  • Jul 22, 2026

    FluidAudio: Run SOTA Audio AI Models Locally on Apple Devices with CoreML

    FluidAudio is a Swift SDK for running state-of-the-art speech-to-text, text-to-speech, speaker diarization, and voice activity detection entirely on-device using Apple's Neural Engine. This post explores its capabilities, architecture, and how to integrate it into your apps.

  • Jul 22, 2026

    SpeechRecognition: A Comprehensive Python Library for Speech-to-Text

    Explore SpeechRecognition, a versatile Python library supporting multiple speech recognition engines and APIs, both online and offline, with practical examples and troubleshooting tips.

  • Jul 11, 2026

    Speech Swift: On-Device AI Speech Toolkit for Apple Silicon

    Explore Speech Swift, an open-source toolkit for ASR, TTS, speech-to-speech, VAD, and diarization on Apple Silicon using MLX and CoreML.

  • Jun 6, 2026

    Miso TTS 8B: A High-Quality Open-Source Text-to-Speech Model

    Miso TTS 8B is a state-of-the-art, open-source text-to-speech model with 8 billion parameters, offering highly emotive speech generation and voice cloning capabilities.

  • May 24, 2026

    Voice-Pro: An Open-Source All-in-One AI Audio & Dubbing Suite

    Voice-Pro is a powerful, open-source Gradio-based WebUI that integrates state-of-the-art voice cloning, transcription, and translation tools into one workflow.

  • May 21, 2026

    OpenLess: The Open-Source AI Voice Input Tool for Developers

    Stop typing, start talking. OpenLess is a cross-platform, privacy-focused tool that turns your voice into structured, AI-polished text directly at your cursor.

  • May 14, 2026

    Supertonic: Lightning-Fast, On-Device Multilingual TTS

    Discover Supertonic, a powerful, open-source text-to-speech system that brings high-quality, multilingual voice synthesis directly to your device. By leveraging ONNX Runtime, Supertonic eliminates the need for cloud APIs, ensuring total privacy and near-instant performance. Whether you are a developer working with Python, C++, Rust, or web technologies, this lightweight engine offers 31-language support and superior reading accuracy for complex text. Learn how this 99M parameter model outperforms larger alternatives in speed and efficiency, making it the perfect choice for edge computing, mobile apps, and browser-based projects. Explore the future of local, private, and lightning-fast speech generation today.

  • Apr 12, 2026

    VoxCPM2: 2B Multilingual TTS with Voice Cloning & Design

    Discover VoxCPM2, the groundbreaking 2B parameter tokenizer-free TTS model supporting 30 languages with studio-quality 48kHz audio. Create voices from text descriptions, clone any speaker with perfect fidelity, and achieve real-time performance (RTF 0.13 on RTX 4090). Fully open-source under Apache 2.0 with Python API, CLI, web demo, LoRA fine-tuning, and production deployment ready. Outperforms commercial models across major TTS benchmarks.

Previous 1 / 3 Next

Curated AI tools, open source projects, tutorials, and resources for developers building with artificial intelligence.

Terms of Service Privacy Policy © 2026 AIBit-Discover Open Source Projects