Tag: #asr
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3 posts
Choosing Speech Models: Practical Criteria for STT and TTS
Unlike text models, speech models are chosen after the language coverage, audio length constraints, real-time requirement, and diarization need are already fixed. This post lays out the parameters, licenses, language cov
2026-08-12 · 7 min read #ai#huggingface#open-source-llm#speech-to-text#text-to-speechVoice AI & TTS 2026 Deep Dive - ElevenLabs · Cartesia Sonic · OpenAI Voice · Play.HT · Hume · Sesame · Fish Audio · Deepgram Aura
2026 is the year voice AI moved from the STT → LLM → TTS pipeline to full-duplex real-time voice agents as the new norm. ElevenLabs v3 holds the multilingual / emotion TTS throne, Cartesia Sonic 3 hits 75ms TTFW and is t
2026-05-16 · 19 min read #voice-ai#tts#elevenlabs#cartesia#openai-voiceVoice AI in 2026 — ElevenLabs / Cartesia / Sesame / Whisper Turbo / Deepgram / Parakeet Deep Dive
In October 2024 Whisper Large v3 Turbo got 8x faster, Cartesia (built by the Mamba authors) hit sub-90ms TTS, and Sesame from Brendan Iribe launched "voice presence." By May 2026 the TTS, STT, and realtime-agent axes hav
2026-05-15 · 22 min read #voice-ai#tts#stt#asr#elevenlabs