How Voice AI Learned to Know When You're Done Talking
Five decades of turn-taking research on one timeline: conversation analysis, VAD and silence timers, predictive models like VAP, semantic endpointing, and full-duplex speech.
Notes & essays
Technical writing on conversational AI, speech, language models, and the things I build.
Five decades of turn-taking research on one timeline: conversation analysis, VAD and silence timers, predictive models like VAP, semantic endpointing, and full-duplex speech.
We built on-device neural text-to-speech for Flowent because React Native had no good answer — 31 languages, no API keys, no network calls, no per-request bill. Now it is open source. Here is the path from install to shipping, rough edges included.
Flowent — the real-time voice agent for language immersion — is now available to download on iOS and Android.
Two vision-language models, the same 150 silent RAVDESS clips, one question — can a VLM read emotion from a face with the audio stripped out? A head-to-head on classification accuracy and per-class behaviour.