arXiv:2512.06757cs.SDcs.CV2025-12被引 2

统一跨模态对齐框架,提升人脸与语音关联的跨语言验证性能。

XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association

  • 融合显式与隐式对齐机制,联合优化人脸与语音特征嵌入。
  • 在MAV-Celeb数据集上表现优异,支持听觉与非听觉语言场景。
  • 适用于跨模态身份验证,尤其适合多语言环境下的应用。

本文提出XM-ALIGN(统一跨模态嵌入对齐框架),用于ICASSP 2026年FAME挑战赛。该框架结合显式与隐式对齐机制,在“听到”和“未听到”语言条件下均显著提升跨模态验证性能。通过从人脸与语音编码器中提取特征嵌入,并利用共享分类器联合优化,采用均方误差(MSE)作为嵌入对齐损失,确保模态间紧密对齐。训练过程中还引入数据增强策略以提升模型泛化能力。实验结果表明,该方法在MAV-Celeb数据集上表现优越。代码将发布于https://github.com/PunkMale/XM-ALIGN。

原文摘要 · Abstract (English)

This paper introduces our solution, XM-ALIGN (Unified Cross-Modal Embedding Alignment Framework), proposed for the FAME challenge at ICASSP 2026. Our framework combines explicit and implicit alignment mechanisms, significantly improving cross-modal verification performance in both "heard" and "unheard" languages. By extracting feature embeddings from both face and voice encoders and jointly optimizing them using a shared classifier, we employ mean squared error (MSE) as the embedding alignment loss to ensure tight alignment between modalities. Additionally, data augmentation strategies are applied during model training to enhance generalization. Experimental results show that our approach demonstrates superior performance on the MAV-Celeb dataset. The code will be released at https://github.com/PunkMale/XM-ALIGN.

跨模态对齐人脸语音多语言嵌入优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。