通过多模态冲突分析,精准识别视频中犹豫与矛盾情绪。
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
- 将多模态信号分段对齐,用轻量模型捕捉跨模态不一致
- 在525个视频上实现0.6133的宏平均F1,远超零样本基线
- 结合专家标签体系进行大模型推理,适合健康行为研究场景
犹豫与矛盾(A/H)是阻碍健康行为改变的冲突性情感状态。视频级A/H识别困难,因信号来自面部、语音、语言和身体动作等多模态间的不一致,且个体表现差异大。本文提出PRISM-AH框架,将冻结的视觉、音频和文本编码器对齐至短时间窗,输入轻量级流式模型以评分跨模态失谐、预测下一窗口的犹豫惊喜信号、发现行为原型,并基于参与者元数据进行条件化。密集的时间窗标注作为辅助监督目标,决策阈值通过宏平均F1校准。知识引导的大语言模型基于数据集的专家线索分类体系,对结构化证据进行推理,仅在验证性能提升时进行晚期融合。在525个视频的公开测试集上,PRISM-AH达到0.6133的宏平均F1,显著优于报告的零样本基线0.2827。推理增益在验证集到更大测试集间具有可迁移性。
原文摘要 · Abstract (English)
Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognition of A/H at the video level is difficult, since the signal arises from disagreement across and within facial, vocal, linguistic, and bodily modalities, and manifests differently across individuals. The proposed PRISM-AH (Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition), is a framework that treats A/H as a multimodal conflict that unfolds over time. Frozen vision, audio, and text encoders are aligned into short time windows and passed to a lightweight streaming model that scores cross-modal dissonance, predicts each next window to expose a hesitation surprise signal, discovers behaviour prototypes, and is conditioned on participant metadata. Dense window-level annotations supervise the model as an auxiliary objective, and the decision threshold is calibrated for macro F1. A knowledge-guided large language model then reasons over structured evidence using the expert cue taxonomy of the dataset, and its verdict is fused late only when validation performance improves. On the labelled public test partition of 525 videos, PRISM-AH attains a macro F1 of 0.6133, compared to the reported zero-shot baseline of 0.2827. The reasoning gain is validated to transfer from validation to the larger test partition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。