用简单特征与可靠校准提升视频中犹豫和矛盾情绪识别效果
Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video

- 融合语言、语音、视觉及可读的停顿特征,通过可靠性门控机制融合
- 非语言特征中的语义空隙时长生成16维特征,独立性最强,准确率0.718
- 固定阈值的AP加权集成比调参更稳定,测试集表现达0.731
我们针对ABAW 2026 BAH挑战任务,解决短访谈视频中犹豫与矛盾情绪(A/H)的识别问题。系统结合情感专用的文本、音频、视觉表征与少量可读的语言停顿线索,通过称为情感标记融合(AMF)的可靠性门控机制进行融合,并采用固定阈值的AP加权集成。我们提出“去ASR时间”:语音识别器虽删除填充词和停顿,但保留其时间戳,由此构建16个特征,构成最强且最独立的非语言通道(AP 0.718,与其他特征相关性0.11–0.36)。控制实验发现:跨模态冲突设计对BAH无显著帮助;语言是主导模态,情感音频为有效补充;校准比模型结构更重要。在小验证集上调参导致过拟合(验证集宏F1 0.741,测试集仅0.690),而固定阈值的AP加权集成在测试集达到0.731。
原文摘要 · Abstract (English)
We address ambivalence and hesitancy (A/H) recognition in the ABAW 2026 BAH Challenge: given a short interview video, predict whether the person shows signs of A/H. Our system combines affect-specialised text, audio, and visual representations with a small set of readable linguistic hesitation cues, fused by a reliability gate we call Affective Marker Fusion (AMF), and finished with a simple AP-weighted ensemble at a fixed decision threshold. We also introduce \emph{ASR-erased time}: speech recognisers delete fillers and hesitation pauses from the transcript, but the chunk timestamps keep the time those events took, and sixteen features built from these gaps form the strongest and most independent non-verbal channel we measured (AP $0.718$, correlation $0.11$--$0.36$ with all other members). Across controlled experiments we find three things: cross-modal conflict design does not reliably help on BAH; language is by far the strongest channel while affect-specialised audio is a useful second; and calibration matters more than architecture. Fitting ensemble weights and a threshold on the small validation split overfits: it scores $0.741$ macro-F1 on validation but only $0.690$ on the untouched test set. AP-weighting at a fixed threshold instead reaches $\mathbf{0.731}$ on test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。