通过最大化类别分离增强人脸与声音的关联效果
Face-Voice Association with Inductive Bias for Maximum Class Separation
- 引入最大类别分离作为归纳偏置,优化跨模态表示
- 在两个任务上达到当前最优性能,提升显著
- 适合关注多模态学习新范式的研究人员
人脸-声音关联是多模态学习中的重要课题,通常通过使同一人的面部和声音嵌入向量接近、不同人之间则分离来实现。以往方法依赖损失函数,而近期分类研究发现,将最大类别分离作为归纳偏置可增强嵌入的判别能力。本文首次将其引入人脸-声音关联领域,提出一种以最大类别分离为归纳偏置的方法,强化不同说话人多模态表示间的区分度。定量实验表明,该方法在两种任务设定下均达到当前最优(SOTA)性能。消融实验进一步显示,结合类间正交性损失时,归纳偏置效果最佳。据我们所知,这是首个在多模态学习中应用并验证最大类别分离归纳偏置有效性的研究,为建立新范式铺平道路。
原文摘要 · Abstract (English)
Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated from those of others. Previous work achieved this with loss functions. Recent advancements in classification have shown that the discriminative ability of embeddings can be strengthened by imposing maximum class separation as inductive bias. This technique has never been used in the domain of face-voice association, and this work aims at filling this gap. More specifically, we develop a method for face-voice association that imposes maximum class separation among multimodal representations of different speakers as an inductive bias. Through quantitative experiments we demonstrate the effectiveness of our approach, showing that it achieves SOTA performance on two task formulation of face-voice association. Furthermore, we carry out an ablation study to show that imposing inductive bias is most effective when combined with losses for inter-class orthogonality. To the best of our knowledge, this work is the first that applies and demonstrates the effectiveness of maximum class separation as an inductive bias in multimodal learning; it hence paves the way to establish a new paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。