arXiv:2603.14976cs.MMcs.CV2026-03被引 1

用文本锚定情感模仿强度,提升噪声环境下的情绪估计鲁棒性。

Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation

  • 以文本为锚点,通过双交叉注意力对齐噪声视觉与音频流。
  • 在六维情绪维度上达到最高皮尔逊相关系数,优于现有方法。
  • 适合真实场景下缺失数据或信号干扰的情绪分析任务。

在自然环境中估算情感模仿强度(EMI)是情感计算中的关键挑战,主要难点在于建模跨异构模态的复杂非线性时序动态,尤其当物理信号受损或缺失时。为此,我们提出TAEMI(Text-Anchored Emotional Mimicry Intensity estimation),用于第10届ABAW竞赛。基于视觉与声学信号易受瞬时环境噪声影响的观察,我们打破传统对称融合范式,利用本质上具有稳定、时间无关语义先验的文本转录作为中心锚点。具体地,引入文本锚定双交叉注意力机制,用稳健的文本查询主动过滤帧级冗余并对齐噪声物理流。此外,为防止真实世界无约束场景中不可避免的数据缺失导致性能崩溃,训练时集成可学习的缺失模态标记与模态丢弃策略。在Hume-Vidmimic2数据集上的大量实验表明,TAEMI能有效捕捉细微情绪变化,并在不完美条件下保持预测鲁棒性。该框架在六个连续情绪维度上实现了当前最优的平均皮尔逊相关系数,显著优于现有基线方法。

原文摘要 · Abstract (English)

Estimating Emotional Mimicry Intensity (EMI) in naturalistic environments is a critical yet challenging task in affective computing. The primary difficulty lies in effectively modeling the complex, nonlinear temporal dynamics across highly heterogeneous modalities, especially when physical signals are corrupted or missing. To tackle this, we propose TAEMI (Text-Anchored Emotional Mimicry Intensity estimation), a novel multimodal framework designed for the 10th ABAW Competition. Motivated by the observation that continuous visual and acoustic signals are highly susceptible to transient environmental noise, we break the traditional symmetric fusion paradigm. Instead, we leverage textual transcript--which inherently encode a stable, time-independent semantic prior--as central anchors. Specifically, we introduce a Text-Anchored Dual Cross-Attention mechanism that utilizes these robust textual queries to actively filter out frame-level redundancies and align the noisy physical streams. Furthermore, to prevent catastrophic performance degradation caused by inevitably missing data in unconstrained real-world scenarios, we integrate Learnable Missing-Modality Tokens and a Modality Dropout strategy during training. Extensive experiments on the Hume-Vidmimic2 dataset demonstrate that TAEMI effectively captures fine-grained emotional variations and maintains robust predictive resilience under imperfect conditions. Our framework achieves a state-of-the-art mean Pearson correlation coefficient across six continuous emotional dimensions, significantly outperforming existing baseline methods.

情绪识别多模态融合鲁棒性文本锚定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。