通过精准对齐与门控融合提升人脸语音关联精度
PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
- 引入正交约束对齐人脸与语音嵌入空间
- 在VoxCeleb上准确率显著优于现有方法
- 适合多模态身份识别与跨模态检索场景
我们研究人脸与语音的关联学习任务,该任务近年来在多模态领域受到关注。现有方法依赖于精心设计的负样本挖掘和远距离边界参数,存在局限性。为此,我们提出在联合嵌入空间中施加正交性约束,以对齐人脸与语音嵌入特征。由于两者嵌入空间特性不同,需先进行精确对齐再融合。本文提出一种增强门控融合机制,有效提升了人脸-语音关联性能。在VoxCeleb数据集上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the distant margin parameter. These issues are addressed by learning a joint embedding space in which orthogonality constraints are applied to the fused embeddings of faces and voices. However, embedding spaces of faces and voices possess different characteristics and require spaces to be aligned before fusing them. To this end, we propose a method that accurately aligns the embedding spaces and fuses them with an enhanced gated fusion thereby improving the performance of face-voice association. Extensive experiments on the VoxCeleb dataset reveals the merits of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。