arXiv:2607.07294cs.ROcs.AI2026-07中稿 · presentation at th…

用多模态信息预测对话轮换,让机器人更懂何时该接话。

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

论文配图:Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
图 1 · 摘自论文原文
  • 融合音视频输入,通过自监督方式预测未来发言
  • 在NoXi数据集上优于现有基线,尤其对复杂轮换事件有效
  • 适合用于调解型人机交互的社交机器人

对话轮换预测是社交机器人参与人-人互动的关键需求,尤其在中介场景中,机器人需预判对话动态而非仅响应停顿。本文提出多模态语音活动投影(MM-VAP)框架,将原有的仅音频VAP方法扩展至同步音视频输入,同时保持自监督未来预测目标。该方法基于原本为语音任务优化的音视频预训练主干网络,通过低秩适应(LoRA)适配到多模态轮换任务。在独立编码说话人后,引入跨说话人注意力机制建模关系动态以预测未来语音活动。此外,设计语义一致性损失,对256维输出空间进行正则化,使其符合高层对话行为模式。在NoXi和NoXi+J数据集上的实验表明,该方法优于当前基线,尤其在部分轮换事件上表现更优。在Haru EDR语料库上的额外评估进一步验证了该方向在面向中介的人机交互中的适用性。

原文摘要 · Abstract (English)

Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective. The proposed approach builds on pretrained audio-visual backbones originally optimized for speech-related tasks and adapts them through Low-Rank Adaptation to the multimodal turn-taking problem. After independent speaker encoding, an inter-speaker attention stage models the relational dynamics required to project future voice activity. In addition, a semantic consistency loss is introduced to regularize the 256-state output space according to higher-level dialogue activity patterns. Experiments on NoXi and NoXi+J showed improvements over the current baselines, particularly for some turn-taking events. Additional evaluation on the Haru EDR corpus further supported the suitability of this direction for mediation-oriented human-robot interaction.

对话预测多模态社交机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。