让对话模型同时听声看人,提升嘈杂环境下的对话连贯性。
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
- 融合音视频信息追踪说话人,改进语音识别与发言时机判断。
- 在干扰环境下错误率降低,对话质量与自然度显著提升。
- 适合开发真实场景中鲁棒的智能语音交互系统。
对话模型在噪声多、多人说话的环境中表现不佳,常生成无关回应并出现不自然的发言切换。本文提出 AV-Dialog,首个结合音视频输入的多模态对话框架,通过音频与视觉线索实现目标说话人追踪、发言时机预测和语义连贯回应生成。该模型采用声学分词与多任务、多阶段训练,在单人、合成及真实音视频对话数据集上表现优异,实现了稳健的流式转录、语义驱动的发言边界检测与准确回复生成,显著提升对话自然度。实验表明,在干扰条件下,其转录错误率更低,发言时机预测更准,且人类评估的对话质量更高。结果证明,‘视听结合’对构建具备说话人感知能力的对话系统至关重要,为实际复杂环境中的语音对话代理提供了新路径。
原文摘要 · Abstract (English)
Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses. By combining acoustic tokenization with multi-task, multi-stage training on monadic, synthetic, and real audio-visual dialogue datasets, AV-Dialog achieves robust streaming transcription, semantically grounded turn-boundary detection and accurate responses, resulting in a natural conversational flow. Experiments show that AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction, and enhancing human-rated dialogue quality. These results highlight the power of seeing as well as hearing for speaker-aware interaction, paving the way for {spoken} dialogue agents that perform {robustly} in real-world, noisy environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。