arXiv:2409.00562eess.AScs.CV2024-09中稿 · the ICNLSP2024 con…被引 11

对比三种音视频融合方法,发现声纹与人脸特征结合效果最佳。

Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification

  • 用声纹和人脸特征分别提取特征,再进行多模态融合。
  • 融合声纹与人脸特征在识别任务中准确率达98.37%。
  • 适合需要高精度身份验证的应用场景。

多模态学习通过整合不同模态信息提升学习效果。本文在语音和人脸两个模态上比较了三种融合策略用于人员识别与验证。采用一维卷积神经网络从语音中提取x-vector,利用预训练的VGGFace2网络和迁移学习处理人脸模态;同时,使用gammatonegram作为语音表征,并结合预训练的Darknet19网络。在VoxCeleb2数据集的118名说话人测试集上,通过K折交叉验证评估单模态与三种提出的多模态策略。结果表明,在识别任务中,gammatonegram与人脸特征的特征融合策略表现最优,准确率达98.37%;在验证任务中,人脸特征与x-vector拼接方式取得0.62%的EER(等错误率)。

原文摘要 · Abstract (English)

Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice and face. In this paper, a one-dimensional convolutional neural network is employed for x-vector extraction from voice, while the pre-trained VGGFace2 network and transfer learning are utilized for face modality. In addition, gammatonegram is used as speech representation in engagement with the Darknet19 pre-trained network. The proposed systems are evaluated using the K-fold cross-validation technique on the 118 speakers of the test set of the VoxCeleb2 dataset. The comparative evaluations are done for single-modality and three proposed multimodal strategies in equal situations. Results demonstrate that the feature fusion strategy of gammatonegram and facial features achieves the highest performance, with an accuracy of 98.37% in the person identification task. However, concatenating facial features with the x-vector reaches 0.62% for EER in verification tasks.

音视频融合身份验证多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。