arXiv:2503.16432cs.HCcs.AI2025-03被引 2

多模态模型让虚拟队友更懂何时该说话,提升对话自然度。

Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay

  • 用文本、视觉、音频和游戏数据融合建模,实时预测发言时机。
  • 准确率达87.3%,宏平均F1为83.0%,优于基线模型。
  • 实测显示对话更流畅,且不增加说话频率,适合跨文化协作场景。

本研究探讨在合作游戏环境中人机交互(HAI)中的多模态发言权预测问题。通过模型构建与用户实验相结合的方式,旨在优化语音对话系统(SDS)的对话动态。模型阶段提出一种基于交叉模态Transformer的深度学习架构,同时融合文本、视觉、音频及游戏上下文数据,实现发言权事件的实时预测。相比基线模型,该模型达到87.3%的准确率与83.0%的宏平均F1分数。随后开展用户研究,以《Don't Starve Together》为场景,对比无发言权预测(n=20)与部署该模型(n=40)两种条件下的交互表现,参与者包含英语与韩语使用者。分析涵盖话语数量、打断频率及对虚拟角色的感知。结果表明,该模型显著提升对话流畅性与自然感,维持均衡互动节奏,未显著改变对话频率。研究揭示了发言权预测对用户体验与交互质量的影响,证明多模态自适应对话代理的潜力。

原文摘要 · Abstract (English)

This study investigates multimodal turn-taking prediction within human-agent interactions (HAI), particularly focusing on cooperative gaming environments. It comprises both model development and subsequent user study, aiming to refine our understanding and improve conversational dynamics in spoken dialogue systems (SDSs). For the modeling phase, we introduce a novel transformer-based deep learning (DL) model that simultaneously integrates multiple modalities - text, vision, audio, and contextual in-game data to predict turn-taking events in real-time. Our model employs a Crossmodal Transformer architecture to effectively fuse information from these diverse modalities, enabling more comprehensive turn-taking predictions. The model demonstrates superior performance compared to baseline models, achieving 87.3% accuracy and 83.0% macro F1 score. A human user study was then conducted to empirically evaluate the turn-taking DL model in an interactive scenario with a virtual avatar while playing the game "Dont Starve Together", comparing a control condition without turn-taking prediction (n=20) to an experimental condition with our model deployed (n=40). Both conditions included a mix of English and Korean speakers, since turn-taking cues are known to vary by culture. We then analyzed the interaction quality, examining aspects such as utterance counts, interruption frequency, and participant perceptions of the avatar. Results from the user study suggest that our multimodal turn-taking model not only enhances the fluidity and naturalness of human-agent conversations, but also maintains a balanced conversational dynamic without significantly altering dialogue frequency. The study provides in-depth insights into the influence of turn-taking abilities on user perceptions and interaction quality, underscoring the potential for more contextually adaptive and responsive conversational agents.

多模态对话系统人机协作虚拟角色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。