视觉线索能提升嘈杂环境下语音交接预测的鲁棒性
Visual Cues Support Robust Turn-taking Prediction in Noise
- 融合视觉信息的多模态模型利用面部线索增强预测
- 在10 dB音乐噪声下准确率达72%,显著优于纯音频模型
- 适合部署在真实嘈杂环境中的人机交互系统
精准的语音交接预测模型(PTTM)对自然人机交互至关重要,但其在噪声环境下的表现仍不明确。本研究考察了模型在实际部署中可能遇到的多种噪声条件下的性能。分析显示,现有模型对噪声极为敏感:在干净语音下保持率高达84%,而10 dB音乐噪声下骤降至52%。通过引入带噪声训练数据,构建包含视觉特征的多模态模型,在10 dB音乐噪声下达到72%准确率。该模型在所有噪声类型与信噪比下均优于纯音频模型,证明视觉线索的有效性;但对新类型噪声泛化能力有限。此外,训练依赖精确转录,限制了使用语音识别生成转录在非清洁场景的应用。代码已公开供后续研究使用。
原文摘要 · Abstract (English)
Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% in clean speech to just 52% in 10 dB music noise. Training with noisy data enables a multimodal PTTM, which includes visual features to better exploit visual cues, with 72% accuracy in 10 dB music noise. The multimodal PTTM outperforms the audio-only PTTM across all noise types and SNRs, highlighting its ability to exploit visual cues; however, this does not always generalise to new types of noise. Analysis also reveals that successful training relies on accurate transcription, limiting the use of ASR-derived transcriptions to clean conditions. We make code publicly available for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。