arXiv:2607.29310cs.CV2026-07被引 1

通过多专家共识与可靠性门控,提升视频中犹豫与矛盾行为的识别准确率。

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

论文配图:CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition
图 1 · 摘自论文原文
  • 融合文本、语音、视觉等多模态特征,构建15种组合的集成模型。
  • 在不重叠参与者数据集上,最终系统宏平均F1达0.7771,优于基线。
  • 仅当三名专家一致时才修正预测,有效防止误纠,适合情绪识别研究者。

犹豫与矛盾(A/H)是通过语言、语音、面部活动等非语言线索表达的细微行为状态。ABAW11 A/H视频识别挑战要求系统为每个自然对话视频分配二元标签。性能以宏平均F1衡量,确保对两类样本同等重视。我们提出CALM-AH,一个融合文本、声学、视觉及衍生行为统计特征的多模态集成模型。构建了15种非空特征分支组合,每组通过验证集交叉熵选择最优三类分类器,并优化决策阈值以最大化验证集宏平均F1。最终结果采用固定硬投票权重(来自BROTHER)。进一步引入可靠性门控多专家共识(RG-MEC),一种锚定保持的决策级集成方法:以CALM-AH为基础预测,当三个互补校正专家(CALM-AH、AffectGPT、基于GPT的语义验证器)一致支持另一类别时才修正;否则保留原始预测。该设计在保证一致性前提下允许双向修正。在参与者互斥的ABAW11数据集上,CALM-AH取得0.7525的宏平均F1,完整RG-MEC系统达到0.7771。

原文摘要 · Abstract (English)

Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recognition Challenge asks systems to assign a binary A/H label to each naturalistic interview video. Performance is measured using Macro-F1 so that recognition of both A/H and No-A/H samples receives equal importance. We present CALM-AH, a multimodal ensemble that combines textual, acoustic, visual, and derived behavioural-statistical features. We construct 15 non-empty combinations of these feature branches. For each combination, we select the best of three classifier families using validation binary cross-entropy and optimise its decision threshold for validation Macro-F1. The resulting binary decisions are combined using fixed hard-voting weights transferred from BROTHER. We further introduce Reliability-Gated Multi-Expert Consensus(RG-MEC), an anchor-preserving decision-level ensemble that combines an initial prediction with three complementary correction experts: CALM-AH, AffectGPT, and a GPT-based semantic verifier. The initial system provides the default prediction. Its label is overridden only when all three correction experts unanimously support the same alternative class; otherwise, the anchor prediction is retained. This unanimity-gated design limits the influence of isolated expert errors while permitting bidirectional correction when task-specific, multimodal-affective, and semantic-pragmatic evidence are fully consistent. On the participant-disjoint ABAW11 dataset, CALM-AH achieves a Macro-F1 of 0.7525, and the complete RG-MEC system achieves 0.7771.

多模态行为识别集成学习情感计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。