融合视觉线索提升双人对话中的发言权预测准确率
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
- 结合语音与面部表情、头部姿态、视线等视觉信息建模
- 视频会议中发言权预测准确率达84%,优于纯语音模型的79%
- 面部表情对预测性能贡献最大,适合人机交互研究者参考
对话中的发言权转换具有丰富的多模态特征。现有预测发言权模型(PTTMs)主要依赖语音,我们提出MM-VAP,一种融合语音与面部表情、头部姿态、视线等视觉线索的多模态模型。在视频会议交互中,其发言权保持/切换预测准确率达84%,优于最先进的纯音频模型(79%)。不同于以往将所有发言权转换统一处理的方式,我们按发言间隔时长分组分析,发现引入视觉特征后,MM-VAP在所有时长区间均表现更优。消融实验表明,面部表情特征贡献最大。因此我们推断:当对话双方可见时,视觉线索对发言权预测至关重要,必须纳入模型。此外,我们验证了自动语音对齐在电话语音上的适用性。本研究首次系统分析多模态发言权预测模型,讨论了未来方向,并公开全部代码。
原文摘要 · Abstract (English)
Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate naturalistic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with visual cues including facial expression, head pose and gaze. We find that it outperforms the state-of-the-art audio-only in videoconferencing interactions (84% vs. 79% hold/shift prediction accuracy). Unlike prior work which aggregates all holds and shifts, we group by duration of silence between turns. This reveals that through the inclusion of visual features, MM-VAP outperforms a state-of-the-art audio-only turn-taking model across all durations of speaker transitions. We conduct a detailed ablation study, which reveals that facial expression features contribute the most to model performance. Thus, our working hypothesis is that when interlocutors can see one another, visual cues are vital for turn-taking and must therefore be included for accurate turn-taking prediction. We additionally validate the suitability of automatic speech alignment for PTTM training using telephone speech. This work represents the first comprehensive analysis of multimodal PTTMs. We discuss implications for future work and make all code publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。