arXiv:2512.04814cs.SDcs.CV2025-12被引 3

构建共享嵌入空间,实现跨语言人脸与声音关联匹配

Shared Multi-modal Embedding Space for Face-Voice Association

  • 分别提取人脸、声音及年龄性别特征,映射到统一嵌入空间
  • 采用自适应角度损失训练,平均等错误率降至23.99%
  • 适用于多语言环境下的人声-人脸配对任务

FAME 2026挑战包含两个高难度任务:在多语言设置下训练人脸与声音的关联,并在未见过的语言上进行测试。本文方法采用独立的单模态处理流程,分别进行通用人脸与声音特征提取,并额外加入年龄性别特征以支持预测。各单模态特征被投影至共享嵌入空间,并使用自适应角度损失(AAM)进行训练。该方法在FAME 2026挑战中获得第一名,平均等错误率(EER)为23.99%。

原文摘要 · Abstract (English)

The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal processing pipelines with general face and voice feature extraction, complemented by additional age-gender feature extraction to support prediction. The resulting single-modal features are projected into a shared embedding space and trained with an Adaptive Angular Margin (AAM) loss. Our approach achieved first place in the FAME 2026 challenge, with an average Equal-Error Rate (EER) of 23.99%.

跨模态人脸声音嵌入空间多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。