August 14, 2026
Generate stereo 48 kHz music locally with HOT-Step CPP, a feature-rich C++/GGML interface for ACE-Step, audio tools, plugins, and model training.
Inflect v2 delivers complete 24 kHz English TTS in two compact models (3.96M and 9.36M params) with no external vocoder. Explore architecture, evaluation, and CPU deployment.
KittenTTS delivers high-quality text-to-speech in models as small as 25MB, running entirely on CPU. Learn how to use it, its API, and why it matters for edge AI.
CrispASR is a unified C++ speech engine built on ggml, supporting 53 ASR and 51 TTS models with zero Python dependencies. Run Cohere Transcribe, Parakeet, Voxtral, Qwen3, and more from a single CLI.
FluidAudio is a Swift SDK for running state-of-the-art speech-to-text, text-to-speech, speaker diarization, and voice activity detection entirely on-device using Apple's Neural Engine. This post explores its capabilities, architecture, and how to integrate it into your apps.
Explore SpeechRecognition, a versatile Python library supporting multiple speech recognition engines and APIs, both online and offline, with practical examples and troubleshooting tips.
Explore Speech Swift, an open-source toolkit for ASR, TTS, speech-to-speech, VAD, and diarization on Apple Silicon using MLX and CoreML.
Miso TTS 8B is a state-of-the-art, open-source text-to-speech model with 8 billion parameters, offering highly emotive speech generation and voice cloning capabilities.
Voice-Pro is a powerful, open-source Gradio-based WebUI that integrates state-of-the-art voice cloning, transcription, and translation tools into one workflow.
Stop typing, start talking. OpenLess is a cross-platform, privacy-focused tool that turns your voice into structured, AI-polished text directly at your cursor.
Discover Supertonic, a powerful, open-source text-to-speech system that brings high-quality, multilingual voice synthesis directly to your device. By leveraging ONNX Runtime, Supertonic eliminates the need for cloud APIs, ensuring total privacy and near-instant performance. Whether you are a developer working with Python, C++, Rust, or web technologies, this lightweight engine offers 31-language support and superior reading accuracy for complex text. Learn how this 99M parameter model outperforms larger alternatives in speed and efficiency, making it the perfect choice for edge computing, mobile apps, and browser-based projects. Explore the future of local, private, and lightning-fast speech generation today.
Discover VoxCPM2, the groundbreaking 2B parameter tokenizer-free TTS model supporting 30 languages with studio-quality 48kHz audio. Create voices from text descriptions, clone any speaker with perfect fidelity, and achieve real-time performance (RTF 0.13 on RTX 4090). Fully open-source under Apache 2.0 with Python API, CLI, web demo, LoRA fine-tuning, and production deployment ready. Outperforms commercial models across major TTS benchmarks.