arXiv:2607.13976cs.CV2026-07被引 1

通过融合跨模态矛盾信号,提升视频中犹豫与模糊情绪的识别精度

CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition

论文配图:CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition
图 1 · 摘自论文原文
  • 用冻结的视觉、音频、文本模型提取特征,再通过冲突融合模块计算跨模态不一致
  • 在验证集上达到0.7155的宏平均F1,测试集准确率0.7439,表现领先
  • 适合做多模态情感分析或人机交互中情绪理解的研究者参考

在非受限视频中检测犹豫与模糊情绪(AH)极具挑战,因目标信号本身具有模糊性,且通过跨模态细微不一致表达,而非典型情感。本文提出CF-Net,参与第3届AH视频识别挑战赛(ABAW 11th, ECCV 2026),针对BAH数据集。该模型采用冻结的SigLIP2、HuBERT和DistilBERT作为视觉、音频与文本骨干网络,对每名说话人特征进行归一化以减少身份泄露,并通过冲突融合模块显式计算跨模态两两不一致。训练采用置信度加权焦点损失、流形混合与模态丢弃;附加的置信度回归头利用模糊标注稳定对边界样本的学习。在BAH验证集上获得0.7155的宏平均F1,在私有测试集上达0.7364(平均精度0.7439)。

原文摘要 · Abstract (English)

Detecting ambivalence and hesitancy (AH) in unconstrained video is challenging because the target signal is inherently ambiguous and expressed through subtle cross-modal incongruence rather than prototypical affect. We present CF-Net, a deep multimodal network submitted to the 3rd Edition of the AH Video Recognition Challenge (ABAW 11th, ECCV 2026), targeting the BAH dataset. CF-Net encodes visual, audio, and transcript streams with frozen SigLIP2, HuBERT, and DistilBERT backbones, normalises backbone features per speaker to reduce identity leakage, and fuses them via a ConflictFusion module that explicitly computes pairwise cross-modal incongruence. Training combines certainty-weighted focal loss, manifold mixup, and modality dropout; an auxiliary certainty-regression head uses ambiguity annotations to stabilise learning on genuinely borderline samples. CF-Net achieves a Macro F1 of 0.7155 on the BAH validation set and 0.7364 (AP = 0.7439) on the private challenge test set.

情绪识别多模态融合跨模态不一致深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。