根据互动阶段动态调整音视频信号可信度,提升连续情绪预测准确率。
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
- 按互动阶段自适应估计音视频可靠性,动态调整融合权重
- 在Aff-Wild2上显著提升一致性相关系数,优于现有融合方法
- 适合应对真实场景中信号不可靠、遮挡等复杂情况
现实环境中连续效价-唤醒度估计面临模态可靠性不一致和音频-视觉信号交互依赖性变化的挑战。现有方法多关注时序建模,却忽视了模态可靠性在不同互动阶段存在显著差异。为此,我们提出SAGE框架,通过显式估计并校准多模态融合中的模态级置信度,实现可靠性感知的融合机制。该机制根据各阶段信号的表征能力动态重平衡音视频表示,防止不可靠信号主导预测。通过将可靠性估计与特征表示分离,所提框架在跨模态噪声、遮挡及交互条件变化下实现更稳定的表情估计。在Aff-Wild2基准上的大量实验表明,SAGE相较于现有融合方法持续提升了一致性相关系数,验证了可靠性驱动建模在连续情感预测中的有效性。
原文摘要 · Abstract (English)
Continuous valence-arousal estimation in real-world environments is challenging due to inconsistent modality reliability and interaction-dependent variability in audio-visual signals. Existing approaches primarily focus on modeling temporal dynamics, often overlooking the fact that modality reliability can vary substantially across interaction stages. To address this issue, we propose SAGE, a Stage-Adaptive reliability modeling framework that explicitly estimates and calibrates modality-wise confidence during multimodal integration. SAGE introduces a reliability-aware fusion mechanism that dynamically rebalances audio and visual representations according to their stage-dependent informativeness, preventing unreliable signals from dominating the prediction process. By separating reliability estimation from feature representation, the proposed framework enables more stable emotion estimation under cross-modal noise, occlusion, and varying interaction conditions. Extensive experiments on the Aff-Wild2 benchmark demonstrate that SAGE consistently improves concordance correlation coefficient scores compared with existing multimodal fusion approaches, highlighting the effectiveness of reliability-driven modeling for continuous affect prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。