Tag: #voice-ai
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 4 posts
Voice AI & TTS 2026 Deep Dive - ElevenLabs · Cartesia Sonic · OpenAI Voice · Play.HT · Hume · Sesame · Fish Audio · Deepgram Aura
2026 is the year voice AI moved from the STT → LLM → TTS pipeline to full-duplex real-time voice agents as the new norm. ElevenLabs v3 holds the multilingual / emotion TTS throne, Cartesia Sonic 3 hits 75ms TTFW and is t
2026-05-16 · 19 min read #voice-ai#tts#elevenlabs#cartesia#openai-voiceVoice AI in 2026 — ElevenLabs / Cartesia / Sesame / Whisper Turbo / Deepgram / Parakeet Deep Dive
In October 2024 Whisper Large v3 Turbo got 8x faster, Cartesia (built by the Mamba authors) hit sub-90ms TTS, and Sesame from Brendan Iribe launched "voice presence." By May 2026 the TTS, STT, and realtime-agent axes hav
2026-05-15 · 22 min read #voice-ai#tts#stt#asr#elevenlabsWebRTC Media Infrastructure 2026 — LiveKit·Pion·Daily·100ms·mediasoup·Janus·Cloudflare Realtime and WHIP/WHEP Deep Dive
Real-time voice and video in 2026 is no longer the 'write WebRTC by hand' game. The explosion of AI voice agents pushed LiveKit into de facto standard status, and Pion settled in as the engine of the Go camp. Daily, 100m
2026-05-14 · 26 min read #webrtc#livekit#pion#mediasoup#dailyVoice AI in Practice — The Complete Guide: Real-Time STT/TTS, Speech LLMs, Turn-Taking, and Deepfake Defense (2025)
Why "AI without a screen" became the hottest product category of 2025. Real-time voice pipelines (VAD/STT/LLM/TTS), speech LLMs (GPT-4o realtime/Gemini Live/Moshi), turn-taking and interruption, emotion and prosody contr
2026-04-15 · 12 min read #voice-ai#stt#tts#gpt-4o-realtime#moshi