arXiv:2512.14961cs.CVcs.SD2025-12被引 1

融合肢体动作、人脸和语音,实现缺失模态下的鲁棒人物识别

Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities

  • 采用多任务学习与交叉注意力融合模态特征,提升信息表达能力
  • 在CANDOR数据集上达99.51%准确率,单模态下仍保持高精度
  • 适合真实场景中模态缺失的人员识别应用

人物识别系统通常依赖音频、视觉或行为线索,但现实环境中常出现模态缺失或退化。为此,本文提出一种融合上半身运动、人脸和语音的多模态人物识别框架。实验表明,身体动作在单会话评估中优于传统的人脸和语音模态,且在多会话场景中作为互补线索提升整体性能。模型采用统一的混合融合策略,同时融合特征级与评分级信息,最大化表征丰富性与决策准确性。具体通过多任务学习独立处理各模态,再利用交叉注意力与门控融合机制挖掘单模态信息与跨模态交互。最后结合置信度加权与错误修正机制,动态适应缺失数据,使单一分类头在单模态与双模态场景下均达最优表现。我们在新提出的基于访谈的多模态数据集CANDOR上进行评估,并首次对其进行基准测试。结果表明,所提三模态系统在人物识别任务中达到99.51%的Top-1准确率。此外,在广泛使用的VoxCeleb1数据集上,双模态模式下准确率达99.92%,超越传统方法。即使缺失一或两个模态,系统仍保持高准确率,展现出强鲁棒性。代码与数据已公开。

原文摘要 · Abstract (English)

Person identification systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently present with missing or degraded modalities. To address this challenge, we propose a multimodal person identification framework incorporating upper-body motion, face, and voice. Experimental results demonstrate that body motion outperforms traditional modalities such as face and voice in within-session evaluations, while serving as a complementary cue that enhances performance in multi-session scenarios. Our model employs a unified hybrid fusion strategy, fusing both feature-level and score-level information to maximize representational richness and decision accuracy. Specifically, it leverages multi-task learning to process modalities independently, followed by cross-attention and gated fusion mechanisms to exploit both unimodal information and cross-modal interactions. Finally, a confidence-weighted strategy and mistake-correction mechanism dynamically adapt to missing data, ensuring that our single classification head achieves optimal performance even in unimodal and bimodal scenarios. We evaluate our method on CANDOR, a newly introduced interview-based multimodal dataset, which we benchmark in this work for the first time. Our results demonstrate that the proposed trimodal system achieves 99.51% Top-1 accuracy on person identification tasks. In addition, we evaluate our model on the VoxCeleb1 dataset as a widely used evaluation protocol and reach 99.92% accuracy in bimodal mode, outperforming conventional approaches. Moreover, we show that our system maintains high accuracy even when one or two modalities are unavailable, making it a robust solution for real-world person recognition applications. The code and data for this work are publicly available.

人物识别多模态鲁棒性动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。