arXiv:2501.03190cs.LGcs.HC2025-01被引 2

用多模态数据预测视频会议中的不流畅与不愉悦时刻

Multimodal Machine Learning Can Predict Videoconference Fluidity and Enjoyment

  • 融合音频、面部动作和身体动作特征进行建模
  • 最高预测准确率达ROC-AUC 0.87,音频特征最关键
  • 适合研究远程沟通体验优化或人机交互的学者

视频会议已成为工作与非正式交流的常见方式,但常缺乏面对面交谈的流畅性与愉悦感。本研究利用多模态机器学习预测视频会议中的负面体验时刻。从RoomReader语料库中采样数千段短片段,提取音频嵌入、面部动作及身体运动特征,训练模型以识别低对话流畅度、低愉悦度,以及分类对话事件(附和、打断或停顿)。最佳模型在独立测试会话中达到最高ROC-AUC 0.87,且领域通用音频特征最为关键。结果表明,多模态音视频信号可有效预测高层次主观对话结果。此外,该研究为视频会议用户体验研究提供了新方法,展示了通过多模态学习识别罕见负面体验时刻的可能性,便于后续分析与干预。

原文摘要 · Abstract (English)

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict moments of negative experience in videoconferencing. We sampled thousands of short clips from the RoomReader corpus, extracting audio embeddings, facial actions, and body motion features to train models for identifying low conversational fluidity, low enjoyment, and classifying conversational events (backchanneling, interruption, or gap). Our best models achieved an ROC-AUC of up to 0.87 on hold-out videoconference sessions, with domain-general audio features proving most critical. This work demonstrates that multimodal audio-video signals can effectively predict high-level subjective conversational outcomes. In addition, this is a contribution to research on videoconferencing user experience by showing that multimodal machine learning can be used to identify rare moments of negative user experience for further study or mitigation.

视频会议多模态学习用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。