Tag: #audio-language-model
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 2 posts
Speech Recognition and Synthesis Technical Reports: What to Read, and What a Single WER Hides
Ten speech recognition and synthesis technical reports, each verified by opening the arXiv abstract page directly. Whisper, Omnilingual ASR, Qwen3-ASR, the Open ASR Leaderboard, Seed-TTS and F5-TTS, CosyVoice 2, Qwen3-TT
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#speech-recognition#text-to-speechSOTA Music and Audio Generation — Neural Codecs and Generative Models
A lineage-focused overview from audio representations (waveform, spectrogram, neural codec) to autoregressive audio language models, diffusion-based audio, and text-to-music conditioning. We analyze the principles of the
2026-06-30 · 9 min read #ai-papers#audio-generation#music-generation#neural-codec#audio-language-model