用语义手势提升对话轮换预测准确率
Modeling Turn-Taking with Semantically Informed Gestures
- 构建融合语义手势的多模态对话数据集
- 加入手势信息后模型性能持续提升
- 适合研究人机交互与多模态对话系统者
对话中人类通过语音、手势和视线等多模态线索管理发言权交接。尽管语言和声学特征已具信息量,手势仍提供互补线索。为此,我们引入 DnD Gesture++,即在多人对话数据集 DnD Gesture 基础上扩展的语义手势标注版本,包含 2,663 条涵盖象形、隐喻、指示及话语类型的手势标注。基于该数据集,我们采用融合文本、音频与手势的 Mixture-of-Experts 框架建模发言权预测。实验表明,引入语义引导的手势可稳定提升基线模型表现,验证了其在多模态发言权管理中的补充作用。
原文摘要 · Abstract (English)
In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide complementary cues for modeling these transitions. To study this, we introduce DnD Gesture++, an extension of the multi-party DnD Gesture corpus enriched with 2,663 semantic gesture annotations spanning iconic, metaphoric, deictic, and discourse types. Using this dataset, we model turn-taking prediction through a Mixture-of-Experts framework integrating text, audio, and gestures. Experiments show that incorporating semantically guided gestures yields consistent performance gains over baselines, demonstrating their complementary role in multimodal turn-taking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。