用手势数据提升语言模型对口语对话的建模能力
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
- 将3D手势序列转为离散符号,与文本嵌入对齐
- 在三个话语线索预测任务中准确率均提升
- 适合研究语音对话与非语言线索融合的学者
语言学研究表明,非语言线索(如手势)在口语对话中起关键作用。例如,说话人通过手势指示话题转换,帮助听者识别语篇转折。本文探究联合建模人体运动序列与语言是否能提升语言模型对口语对话的建模效果。我们首先使用VQ-VAE将3D人体运动序列编码为离散手势符号,再通过特征对齐将其映射到文本嵌入空间。为评估该模型在口语对话中的表现,构建了针对三类话语线索的文本补全任务:话语连接词、立场标记和量化词。结果表明,融入手势信息后,三项任务的标记预测准确率均有所提升,验证了手势提供的互补性信息对口语建模的有效性。本工作是利用非语言线索改进语言模型口语建模的初步探索。
原文摘要 · Abstract (English)
Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse. For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse. In this work, we investigate whether the joint modeling of gestures using human motion sequences and language can improve spoken discourse modeling in language models. To integrate gestures into language models, we first encode 3D human motion sequences into discrete gesture tokens using a VQ-VAE. These gesture token embeddings are then aligned with text embeddings through feature alignment, mapping them into the text embedding space. To evaluate the gesture-aligned language model on spoken discourse, we construct text infilling tasks targeting three key discourse cues grounded in linguistic research: discourse connectives, stance markers, and quantifiers. Results show that incorporating gestures enhances marker prediction accuracy across the three tasks, highlighting the complementary information that gestures can offer in modeling spoken discourse. We view this work as an initial step toward leveraging non-verbal cues to advance spoken language modeling in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。