融合语言、声音和视觉信号,提升人机对话中的轮换与回应预测准确率。
Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals
- 利用多模态信号构建端到端预测框架,支持任意组合输入。
- 在1.5百万词数据上实现转接与回应预测性能显著提升。
- 公开210小时视频数据集,助力后续对话系统研究。
本文针对人机对话中轮换与回应行为预测的不足,提出一种基于多模态信号(语言、声学、视觉)的解决方案。为克服现有数据集局限,我们设计了一套自动数据采集流程,收集并标注了超过210小时的人类对话视频。基于此构建了多模态面对面(MM-F2F)对话数据集,包含约150万词及对应约2000万帧的轮换与回应标注。同时,提出一个端到端框架,从多模态信号中预测轮换与回应概率。模型强调模态间关联性,可灵活适配文本、音频、视频任意组合,适用于多种真实场景。实验表明,该方法在轮换预测上F1得分提升10%,在回应预测上提升33%,达到当前最优水平。数据集与代码已公开,便于后续研究。
原文摘要 · Abstract (English)
This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an automatic data collection pipeline that allows us to collect and annotate over 210 hours of human conversation videos. From this, we construct a Multi-Modal Face-to-Face (MM-F2F) human conversation dataset, including over 1.5M words and corresponding turn-taking and backchannel annotations from approximately 20M frames. Additionally, we present an end-to-end framework that predicts the probability of turn-taking and backchannel actions from multi-modal signals. The proposed model emphasizes the interrelation between modalities and supports any combination of text, audio, and video inputs, making it adaptable to a variety of realistic scenarios. Our experiments show that our approach achieves state-of-the-art performance on turn-taking and backchannel prediction tasks, achieving a 10% increase in F1-score on turn-taking and a 33% increase on backchannel prediction. Our dataset and code are publicly available online to ease of subsequent research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。