arXiv:2409.09135cs.AIcs.CL2024-09被引 13

用大模型融合多模态数据预测对话中参与度,提升人机交互理解力。

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

  • 通过大模型将语音与非语言行为融合成多模态文本进行推理
  • 在34人对话数据上达到现有方法相当的预测性能
  • 适合研究人机交互、心理健康与无障碍沟通的学者

过去十年,可穿戴设备(如智能眼镜)在传感器技术、设计与处理能力上取得显著进展,为高密度人类行为数据采集带来新机遇。配备摄像头的智能眼镜可无感记录自然情境下的非语言行为,帮助分析双人互动中的参与度。本文聚焦于通过分析言语与非言语线索预测互动参与度,以识别冷漠或困惑信号。研究收集了34名参与者在轻松对话中的数据,并获取其事后自评的参与度评分。提出一种基于大语言模型(LLMs)的新融合策略,将多模态行为信息整合为「多模态转录本」,供大模型进行行为推理。初步实验表明,该方法性能已接近传统融合技术,展现出巨大优化潜力。这是首个尝试通过语言模型对真实人类行为进行「推理」的研究之一。所采集的特征与数据将公开,推动后续研究。智能眼镜使我们能持续获取高密度多模态行为数据,为改善人类沟通提供新路径,具有重要社会价值。

原文摘要 · Abstract (English)

Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new era of opportunity for high-density human behavior data. Equipped with wearable cameras, these glasses offer a unique opportunity to analyze non-verbal behavior in natural settings as individuals interact. Our focus lies in predicting engagement in dyadic interactions by scrutinizing verbal and non-verbal cues, aiming to detect signs of disinterest or confusion. Leveraging such analyses may revolutionize our understanding of human communication, foster more effective collaboration in professional environments, provide better mental health support through empathetic virtual interactions, and enhance accessibility for those with communication barriers. In this work, we collect a dataset featuring 34 participants engaged in casual dyadic conversations, each providing self-reported engagement ratings at the end of each conversation. We introduce a novel fusion strategy using Large Language Models (LLMs) to integrate multiple behavior modalities into a ``multimodal transcript'' that can be processed by an LLM for behavioral reasoning tasks. Remarkably, this method achieves performance comparable to established fusion techniques even in its preliminary implementation, indicating strong potential for further research and optimization. This fusion method is one of the first to approach ``reasoning'' about real-world human behavior through a language model. Smart glasses provide us the ability to unobtrusively gather high-density multimodal data on human behavior, paving the way for new approaches to understanding and improving human communication with the potential for important societal benefits. The features and data collected during the studies will be made publicly available to promote further research.

多模态融合大模型应用行为预测智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。