arXiv:2603.14361cs.CV2026-03被引 8

通过智能融合多模态数据,提升自然场景中犹豫与矛盾行为的识别准确率。

BROTHER: Behavioral Recognition Optimized Through Heterogeneous Ensemble Regularization for Ambivalence and Hesitancy

  • 设计多模态特征提取与粒子群优化集成策略,抑制过拟合。
  • 在未见测试集上达到0.7465的宏F1分数,优于单一模态。
  • 适合做情感计算、人机交互中的复杂行为分析研究者参考。

在自然视频场景中识别犹豫与矛盾等复杂行为状态仍是情感计算的重大挑战。与基础面部表情不同,这类行为表现为细微的多模态冲突,需深度上下文与时间理解。本文提出一种高度正则化的多模态融合流程,在视频级别预测犹豫与矛盾。从视觉、听觉和语言数据中提取鲁棒单模态特征,引入专门设计的统计文本模态以捕捉语音时序变化与行为线索。通过评估15种模态组合在多层感知机(MLP)、随机森林(Random Forest)与梯度提升树(GBDT)构成的分类器委员会上的表现,基于验证集二元交叉熵(BCE)损失选择最校准模型。为避免过拟合训练分布,采用粒子群优化(PSO)硬投票集成,其适应度函数动态引入训练-验证差距惩罚项(lambda),主动抑制冗余或过拟合分类器。综合评估显示,语言特征是独立预测能力最强的模态,而经过强正则化处理的PSO集成(lambda = 0.2)有效整合多模态协同效应,在未见测试集上取得最高0.7465的宏F1分数。结果表明,将犹豫与矛盾视为多模态冲突,并通过智能加权集成框架评估,可构建鲁棒的野外行为分析系统。

原文摘要 · Abstract (English)

Recognizing complex behavioral states such as Ambivalence and Hesitancy (A/H) in naturalistic video settings remains a significant challenge in affective computing. Unlike basic facial expressions, A/H manifests as subtle, multimodal conflicts that require deep contextual and temporal understanding. In this paper, we propose a highly regularized, multimodal fusion pipeline to predict A/H at the video level. We extract robust unimodal features from visual, acoustic, and linguistic data, introducing a specialized statistical text modality explicitly designed to capture temporal speech variations and behavioral cues. To identify the most effective representations, we evaluate 15 distinct modality combinations across a committee of machine learning classifiers (MLP, Random Forest, and GBDT), selecting the most well-calibrated models based on validation Binary Cross-Entropy (BCE) loss. Furthermore, to optimally fuse these heterogeneous models without overfitting to the training distribution, we implement a Particle Swarm Optimization (PSO) hard-voting ensemble. The PSO fitness function dynamically incorporates a train-validation gap penalty (lambda) to actively suppress redundant or overfitted classifiers. Our comprehensive evaluation demonstrates that while linguistic features serve as the strongest independent predictor of A/H, our heavily regularized PSO ensemble (lambda = 0.2) effectively harnesses multimodal synergies, achieving a peak Macro F1-score of 0.7465 on the unseen test set. These results emphasize that treating ambivalence and hesitancy as a multimodal conflict, evaluated through an intelligently weighted committee, provides a robust framework for in-the-wild behavioral analysis.

行为识别多模态融合情感计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。