研究视频对话中手势与语音如何随对方可见性变化传递信息
Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

- 用语音和骨骼手势联合建模,判断对话中指代对象
- 单独手势就能预测目标,融合模型在语音不确定时效果最好
- 适合研究人机交互、多模态对话的学者参考
情境化语言使用具有多模态和具身特征。例如,手势可携带语音信号中缺失或不明确的信息,但现有对话模型通常仅依赖文本。本文研究在不同对话伙伴可见条件下,手势及其与语音结合所携带的指代信息量。构建了基于语音文本、手势骨架表示或两者融合的模型,用于视频化指代交流任务中识别目标对象。结果表明,仅凭手势即可预测目标,且多模态融合在基于语音的模型不确定时收益最大。仅训练阶段对齐学习表征与目标图像,进一步提升了融合模型性能。与人类交互数据对比发现,对话者可见性影响手势生成与信息量,并存在言语及多模态表现的同步效应,但手势表现无此效应。本研究为人类对话中多模态信息的技术建模与基于训练模型表征的人类交互分析提供了贡献。
原文摘要 · Abstract (English)
Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。