用手势和语言联合解析对话中的指代,提升人机交互自然度
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
- 自监督学习手势表征,将身体动作与语音语义对齐
- 手势特征能显著提升指代消解准确率,即使无语音也可用
- 融合对话历史可进一步提升效果,适合多模态交互研究
在面对面交流中,人们通过语言和手势等多种模态传递信息并解决指代问题。然而,从计算角度研究代表型手势如何指代物体仍不充分。本文提出一种以代表型手势为核心的多模态指代消解任务,并解决手势表征学习的鲁棒性挑战。我们设计了一种自监督预训练方法,使手势表示与语音内容对齐。实验表明,所学嵌入与专家标注高度一致,具备强预测能力。当使用多模态手势表示时,即使推理阶段无语音输入,指代消解准确率仍显著提升;同时利用对话历史也能进一步改善性能。结果表明,手势与语言在指代消解中具有互补作用,为更自然的人机交互建模迈出关键一步。
原文摘要 · Abstract (English)
In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects remains understudied from a computational perspective. In this work, we address this gap by introducing a multimodal reference resolution task centred on representational gestures, while simultaneously tackling the challenge of learning robust gesture embeddings. We propose a self-supervised pre-training approach to gesture representation learning that grounds body movements in spoken language. Our experiments show that the learned embeddings align with expert annotations and have significant predictive power. Moreover, reference resolution accuracy further improves when (1) using multimodal gesture representations, even when speech is unavailable at inference time, and (2) leveraging dialogue history. Overall, our findings highlight the complementary roles of gesture and speech in reference resolution, offering a step towards more naturalistic models of human-machine interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。