arXiv:2607.21347cs.CV2026-07中稿 · publication at IEE…

通过动态表情信息提升人脸识别,尤其在光照差、遮挡等复杂环境下表现更优。

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

论文配图:Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features
图 1 · 摘自论文原文
  • 设计自适应融合机制,根据每段视频质量动态调整音视频贡献权重。
  • 在真实场景数据集上实现0.472的评估相关系数,比基线提升显著。
  • 无需专门训练身份信息,即可有效识别个体,适合弱监督身份验证场景。

传统人脸识别依赖静态外观特征,在表情变化、遮挡和光线不良等非受限环境下性能下降。本文假设音视频表达动态包含补充于静态外观的身份判别信息,且提取该信号需对真实视频中可变输入质量具备鲁棒性的多模态表示。为此,将多模态效价-唤醒度(VA)估计作为预训练任务,提出质量感知自适应融合(QAAF)方法,通过学习每样本每模态的可靠性,并利用可学习软门控与质量依赖丢弃机制动态调整各模态贡献。在Aff-wild2数据集上,基于晚期融合集成的QAAF实现平均一致相关系数(CCC)0.472,优于相同设置下的基线集成(0.415)及单骨干基线(0.288)。此外,当缺失一个模态时,QAAF的CCC仅下降7.5%-34.4%,表现出更强鲁棒性。进一步探究发现,经VA训练的特征虽未进行身份特定训练,但在AFEW-VA(67名演员)和YTF(1,595名受试者)上排名领先于其他软生物特征方法;与ArcFace进行分数级融合后,两数据集误接受率(EER)分别从0.022降至0.021,0.106降至0.104,纠正了AFEW-VA上ArcFace 68.2%的误接受。这些结果确立了多模态VA估计作为补充传统人脸识别的软生物特征新范式。

原文摘要 · Abstract (English)

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video. To learn such representations, we cast multimodal valence-arousal (VA) estimation as a pretext task and propose Quality-Aware Adaptive Fusion (QAAF), which estimates per-sample, per-modality reliability and adapts each modality's contribution through learned soft gating and a quality-dependent dropout. For the problem of VA estimation, QAAF achieves an average Concordance Correlation Coefficient (CCC) of 0.472 via late fusion ensembling on Aff-wild2, improving over a baseline ensemble under the same setting (0.415) as well as a single-backbone baseline (0.288). Furthermore, the proposed QAAF demonstrates greater resilience to unavailable modalities, with only a 7.5-34.4% relative decrease in CCC when one modality is missing. We then probe whether these VA-trained features encode identity without identity-specific training. On AFEW-VA (67 actors) and YTF (1,595 subjects), VA-trained backbone features rank first among evaluated soft biometric methods, and score-level fusion with ArcFace lowers EER on both datasets (0.022 to 0.021 on AFEW-VA, 0.106 to 0.104 on YTF), correcting 68.2% of ArcFace's false accepts on AFEW-VA. These findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition.

多模态融合软生物识别情感计算身份识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。