仅用单麦克风和单摄像头实现多人对话轮换预测,适合真实场景应用。
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

- 通过人脸追踪将语音活动投影到角色关系,实现说话人感知的轮换预测。
- 在双人和三人对话中,对下一说话人预测准确率超越现有基线模型。
- 构建了31小时未剪辑的真实多人对话数据集,支持因果跟踪研究。
当前多说话人轮换预测模型通常依赖复杂的麦克风阵列或多摄像机系统,限制了其在人机交互中的应用。我们提出MuVAP,一种因果多模态框架,通过将语音活动投影与人脸轨迹结合,仅用单声道音频流和单摄像头视角即可实现说话人感知的轮换预测。为解决多说话人建模带来的组合复杂性,提出角色相对投影(Role-Relative Projection),将任意N人交互映射为固定当前发言者与下一位发言者状态。由于现有音视频数据集存在破坏因果追踪的剪辑片段,我们构建了31小时未编辑、单摄像头采集的多人对话语料库。评估显示,MuVAP在双人和三人场景下,于Shift-Hold及下一说话人预测任务中均优于强基线模型。
原文摘要 · Abstract (English)
Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single-camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。