South Korean voice AI pioneers develop conversational systems that respond before users finish speaking

Developers in South Korea are pushing the boundaries of voice AI, creating models capable of real-time, human-like conversation with emotion detection and interruption handling, signalling a new era for intelligent assistants worldwide.

Personal AI services are moving beyond simple question-and-answer tools and into systems that can keep pace with normal conversation. In South Korea, that shift is now visible in voice AI, where developers are focusing on models that can hear interruptions, detect emotion and respond before a user has finished a thought. The result is a new generation of assistants designed to sound less mechanical and to behave more like a person in a live exchange.

Naver Cloud has placed that ambition at the centre of its latest research. According to the company, its new voice AI system, Sommelier, was presented at ACL 2026, one of the most important academic conferences in natural language processing. The model uses full-duplex processing, meaning it can handle voice, video and text at the same time rather than waiting for one speaker to finish before the other responds. In practical terms, that allows the system to listen while it is speaking, recognise a follow-up question or interruption and adjust its answer on the fly. Naver said the approach is intended to support future work assistants, customer service tools and real-time interpretation services.

Kakao is pursuing a similar goal with a different emphasis: giving users direct control over tone, emotion and speaking style. The company has said its updated Kanana-o model can follow natural-language instructions such as speaking softly, answering more slowly or using a particular dialect. Kakao says it developed a language-model-friendly voice tokenizer called LM-SPT to encode acoustic features alongside meaning, and then used online reinforcement learning to improve how closely the model follows speech instructions without making the audio sound unnatural. In materials published by the company, Kanana-o scored 94.50 on the Korean InstructTTSEval benchmark, ahead of GPT-4o mini tts at 91.10. Kakao has also said it wants the system to handle non-verbal sounds such as laughter, sighs and interjections more naturally.

The broader competitive picture extends well beyond South Korea. OpenAI has introduced GPT-Live for ChatGPT, a voice feature built to reduce the stop-start feel of older systems by improving endpoint detection and streaming so users can interrupt, restart or change direction mid-answer. Google has also pushed ahead with Gemini 3.1 Flash Live, which is designed to handle hesitations, interruptions and longer reasoning chains while adjusting its response to a user’s tone. The technology is likely to spread into smart glasses, headphones, smart speakers and other wearables, where voice will often be the main interface. For the industry, the issue is no longer whether AI can answer a question, but whether it can hold a conversation that feels immediate, responsive and genuinely human.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.