arXiv:2511.14749cs.CV2025-11被引 2

用视觉大模型提升视频互动分析中的标注质量,改善噪声标签影响。

Vision Large Language Models Are Good Noise Handlers in Engagement Analysis

  • 利用问卷提取行为线索,区分高/低可信度数据集
  • 结合课程学习与软标签优化,逐步引入模糊样本
  • 在多个基准上超越现有方法,提升1.21%以上

视频数据中的互动识别任务受主观标注和噪声标签严重影响。为应对这一挑战,我们提出一种基于视觉大语言模型(VLM)的框架,用于修正标注并指导训练。该框架通过问卷收集行为线索,将数据划分为高可靠性和低可靠性子集。同时引入结合课程学习与软标签优化的训练策略,逐步融入模糊样本,并根据不确定性调整监督信号。实验表明,使用高质量子集训练的经典计算机视觉模型,在本策略下表现显著提升。该方法在EngageNet(六种特征设置中三组优于前序最佳,最高提升+1.21%)、DREAMS与PAFE等基准上分别取得F1分数+0.22与+0.06的增益。

原文摘要 · Abstract (English)

Engagement recognition in video datasets, unlike traditional image classification tasks, is particularly challenged by subjective labels and noise limiting model performance. To overcome the challenges of subjective and noisy engagement labels, we propose a framework leveraging Vision Large Language Models (VLMs) to refine annotations and guide the training process. Our framework uses a questionnaire to extract behavioral cues and split data into high- and low-reliability subsets. We also introduce a training strategy combining curriculum learning with soft label refinement, gradually incorporating ambiguous samples while adjusting supervision to reflect uncertainty. We demonstrate that classical computer vision models trained on refined high-reliability subsets and enhanced with our curriculum strategy show improvements, highlighting benefits of addressing label subjectivity with VLMs. This method surpasses prior state of the art across engagement benchmarks such as EngageNet (three of six feature settings, maximum improvement of +1.21%), and DREAMS / PAFE with F1 gains of +0.22 / +0.06.

视觉大模型互动分析噪声处理标签优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。