arXiv:2505.19328cs.CVcs.LG2025-05被引 16

首个视频多模态犹豫情绪数据集,助力数字健康干预个性化

BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change

  • 构建1427段视频数据集,涵盖面部、语音等多模态犹豫信号
  • 总时长10.6小时,300名参与者,标注了犹豫出现的时间戳
  • 适合做情绪识别、数字健康、人机交互的研究者使用

犹豫与矛盾(A/H)是个人延迟或放弃健康行为改变的主要原因,表现为正负情感冲突,在面部、语音、肢体语言等多模态间或内部存在不一致。尽管专家可识别此类情绪,但在线干预中应用成本高且效果差。因此,自动识别A/H对个性化数字行为干预至关重要。本文提出首个用于视频多模态识别的A/H数据集BAH,包含1,427段视频,总时长10.60小时,来自加拿大300名参与者,通过预设问题诱发并记录了真实在线干预场景中的情绪反应。数据集由三位专家标注,提供帧级与视频级的时间戳及情绪线索标签。附带视频转录文本、对齐人脸图像和用户元数据。由于A与H在实践中表现相似,采用二值标注表示是否存在A/H。论文还提供了基线模型在帧级与视频级识别任务上的基准结果,性能有限,表明需开发更适配的多模态时空模型。数据与代码已公开。

原文摘要 · Abstract (English)

Ambivalence and hesitancy (A/H), closely related constructs, are the primary reasons why individuals delay, avoid, or abandon health behaviour changes. They are subtle and conflicting emotions that sets a person in a state between positive and negative orientations, or between acceptance and refusal to do something. They manifest as a discord in affect between multiple modalities or within a modality, such as facial and vocal expressions, and body language. Although experts can be trained to recognize A/H as done for in-person interactions, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital behaviour change interventions. However, no datasets currently exist for the design of machine learning models to recognize A/H. This paper introduces the Behavioural Ambivalence/Hesitancy (BAH) dataset collected for multimodal recognition of A/H in videos. It contains 1,427 videos with a total duration of 10.60 hours, captured from 300 participants across Canada, answering predefined questions to elicit A/H. It is intended to mirror real-world digital behaviour change interventions delivered online. BAH is annotated by three experts to provide timestamps that indicate where A/H occurs, and frame- and video-level annotations with A/H cues. Video transcripts, cropped and aligned faces, and participant metadata are also provided. Since A and H manifest similarly in practice, we provide a binary annotation indicating the presence or absence of A/H. Additionally, this paper includes benchmarking results using baseline models on BAH for frame- and video-level recognition, and different learning setups. The limited performance highlights the need for adapted multimodal and spatio-temporal models for A/H recognition. The data and code are publicly available.

情绪识别多模态数字健康数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。