arXiv:2604.11730cs.CVcs.HC2026-04被引 3

用视频识别健康干预中的犹豫情绪,助力个性化数字医疗

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

论文配图:Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
图 1 · 摘自论文原文
  • 融合多模态视频信息,用深度学习自动识别犹豫情绪
  • 在BAH数据集上表现有限,凸显建模跨模态矛盾的难度
  • 适合做数字健康、情感计算和个性化干预的研究者

基于行为科学的健康干预通过提供框架帮助患者养成并维持健康习惯,改善医疗结果。面对面干预成本高且难规模化,尤其在资源匮乏地区。数字健康干预具有成本效益,可支持独立生活与自我管理。通过机器学习自动化此类干预近年来受到广泛关注。犹豫与矛盾(A/H)是导致个体延迟、回避或放弃健康干预的主要原因。这类情绪微妙且矛盾,表现为对行为的正负评价之间摇摆,或接受与拒绝之间的冲突,常体现为多模态(如语言、面部、声调、肢体动作)间或内部的一致性不一致。虽然专家可识别,但融入数字系统成本高且效果有限。因此,自动识别A/H对实现个性化与成本效益至关重要。本文探索深度学习在视频中识别多模态A/H的应用,涵盖三种学习设置:监督学习、无监督域适应以实现个性化,以及通过大语言模型(LLMs)进行零样本推理。实验基于最新发布的、独特的BAH视频数据集。结果表明性能有限,说明需要更适配的多模态模型。更好的时空建模与跨模态融合方法对捕捉模态内/间矛盾至关重要。

原文摘要 · Abstract (English)

Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has recently gained considerable attention. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.

多模态情绪识别数字健康零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。