通过凸包嵌入提升音容关联准确率
Face and Voice Cross-modal Association with Learning Convex Feature Embedding
- 用凸包约束音容特征,使同人音容更靠近
- 在VoxCeleb上验证,误报率显著降低
- 适合跨模态匹配与检索任务
音容关联学习是深度学习中的挑战性任务。本文提出一种简单而强大的跨模态特征嵌入方法,用于人脸与语音的关联。以往工作虽关注跨模态判别,却忽视了音频与视频特征间的异质性,导致大量误报与漏报。为解决该问题,本文方法在跨模态特征间引入另一特征,使同一人的音容特征被嵌入到一个凸包中。结合跨模态注意力机制与凸嵌入技术,有效抑制误报与漏报,通过最小化类间差异实现。在大规模VoxCeleb数据集上对跨模态验证、匹配与检索任务进行了全面评估。实验结果表明,所提方法显著优于现有最先进方法。
原文摘要 · Abstract (English)
Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。