arXiv:2506.13971eess.AScs.CL2025-06被引 1

用少量标注数据,精准识别视频会议中的不流畅时刻

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

  • 融合音视频文本的半监督学习,利用少量标签数据训练模型
  • 仅用8%标注数据,性能达到全量标注模型的96%
  • 适合需要高效标注的会议体验研究或智能会议系统开发

视频会议中的群体对话是复杂的社交行为,但其中主观上感到不流畅或无趣的时刻在自然数据中极为稀少,导致监督学习需大量人工标注。本文采用半监督学习,结合有标签与无标签片段,训练多模态(音频、面部表情、文本)深度特征,以预测会议中非流畅或不愉快的瞬间。基于模态融合的协同训练半监督模型达到ROC-AUC 0.9,F1分数0.6,在相同标注量下比监督模型提升最高4%;尤其值得注意的是,仅使用8%标注数据的最优半监督模型,性能达到监督模型全量数据时的96%。该方法为建模视频会议体验提供了高效的标注策略。

原文摘要 · Abstract (English)

Group conversations over videoconferencing are a complex social behavior. However, the subjective moments of negative experience, where the conversation loses fluidity or enjoyment remain understudied. These moments are infrequent in naturalistic data, and thus training a supervised learning (SL) model requires costly manual data annotation. We applied semi-supervised learning (SSL) to leverage targeted labeled and unlabeled clips for training multimodal (audio, facial, text) deep features to predict non-fluid or unenjoyable moments in holdout videoconference sessions. The modality-fused co-training SSL achieved an ROC-AUC of 0.9 and an F1 score of 0.6, outperforming SL models by up to 4% with the same amount of labeled data. Remarkably, the best SSL model with just 8% labeled data matched 96% of the SL model's full-data performance. This shows an annotation-efficient framework for modeling videoconference experience.

视频会议多模态融合半监督学习标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。