arXiv:2607.12176cs.CV2026-07

通过多模态融合与校准集成,提升对视频中犹豫与矛盾情绪的识别准确率。

HEDGE: A Calibrated Ensemble for A/H Recognition

  • 采用冻结特征的三模型等权重集成,融合人脸、音频、文本和姿态信息。
  • 在公开测试集上达0.7358宏F1,私有测试集最高0.7367,表现稳定可靠。
  • 文本单独建模已接近全系统性能,表明模态冗余是系统鲁棒的关键。

犹豫与矛盾(A/H)削弱数字行为干预效果,从视频中自动识别A/H是BAH数据集上ABAW挑战的目标。我们提出HEDGE(基于分布感知的广义集成的犹豫/矛盾估计),用于第11届挑战赛:一个校准后的等权重集成系统,包含三个融合模型,分别处理冻结的人脸、音频、文本和姿态嵌入,在公开测试集上达到0.7358宏F1。我们还提交了四种变体至今年私有测试(30名新参与者):固定阈值版本、单一模型、以及带缓存个性化测试时适应(TTA)的集成。纯校准集成在私有测试集上得分为0.7361,与公开测试预测高度一致;而TTA变体以0.7367得分位居五项结果之首,为官方挑战成绩(团队AIWELL,排名第5/12),尽管其在公开测试中未见增益。单一模型降至0.6759,远低于所有集成方案。我们解释该现象:仅私有测试存在分布偏移需纠正,而公开测试无此问题;此外,仅文本线性探测已达0.716宏F1(与全系统误差相关性0.91),说明系统鲁棒性源于模态冗余,单一路径则更脆弱。我们还报告超过60项受控实验(模态、主干、损失函数、适应策略消融),均未超越当前系统;可解释性分析显示,语篇表达风格主导信号,非语言信号上限接近0.60宏F1。

原文摘要 · Abstract (English)

Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW A/H challenge on the BAH dataset. We describe HEDGE (Hesitancy/Ambivalence Estimation via Distribution-aware, Generalized Ensembling), our system for the 11th edition of the challenge: a calibrated, equal-weight ensemble of three fusion models over frozen face, audio, text, and pose embeddings, which reaches 0.7358 macro-F1 on the public test set. We also submitted four variants of this system to this year's private test (30 new participants): a fixed-threshold version, a single individual model, and an ensemble with cache-personalization test-time adaptation (TTA). The plain calibrated ensemble scored 0.7361 macro-F1 on the private test, closely matching our public-test estimate, and the TTA variant scored highest of all five at 0.7367 macro-F1, our official challenge result (team AIWELL, rank 5 of 12), even though TTA showed no benefit on the public test. The single individual model dropped to 0.6759, far more than any ensemble variant. We explain both results: TTA only has distribution shift to correct on the private test, which the public test lacks, and a text-only linear probe reaches 0.716 macro-F1 (within noise of the full system, correlated at 0.91 in its errors), so the ensemble's robustness to new participants comes from the same modality redundancy that makes single, less-diversified models comparatively brittle. We additionally report a systematic study of more than 60 further controlled experiments (modality, backbone, loss, and adaptation ablations) that did not improve on this system, and an explainability analysis showing the transcript's delivery style dominates the signal while the extractable non-verbal ceiling saturates near 0.60 macro-F1.

情绪识别多模态融合集成学习行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。