Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
-
Updated
Sep 14, 2026 - Python
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
Turn detection for full-duplex dialogue communication
A real-time software for turn-taking, backchannel, and head-nodding prediction
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
Python package based on Gibson's framework (2003) for turn-taking in group conversation analysis.
Local-first VAD, barge-in, and turn-taking primitives for interruptible voice agents.
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
🏀💬 Advanced User Interfaces course project developed as a 1st year student of MSc in CS Eng. at Politecnico di Milano
Full-duplex conversational turn-taking in C++17 — no ASR, no LLM, no neural VAD. Decides when you have finished speaking from prosody alone, and hears you interrupt it through its own speaker echo.
Open middleware protocol to make voice AI agents sound human — ambient audio, turn-taking, spoken register. Built on Pipecat.
To associate your repository with the turn-taking topic, visit your repo's landing page and select "manage topics."