arXiv:2512.20376cs.CV2025-12中稿 · ICASSP 2026被引 5

跨语言人脸与声音关联,突破语种限制的多语交流新方法

Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge

  • 构建跨语言人脸-声音匹配模型,支持测试时语言不同于训练
  • 在多语言场景下实现超过70%的准确率,优于基线方法
  • 适合多语种语音识别、跨模态身份验证等应用

全球超过一半人口为双语者,多语言沟通场景日益普遍。2026年国际声学、语音与信号处理会议(ICASSP 2026)举办的跨语言人脸-声音关联挑战赛(FAME 2026),旨在开发在测试阶段语言与训练阶段不一致时仍有效的脸-声关联方法。本报告简要总结了该挑战赛的背景、任务设定及主要成果。参赛方案聚焦于建模跨语言特征不变性,通过多模态对齐机制,在不同语言组合下实现稳定匹配性能。实验表明,最优方法在多种跨语言设置中达到70%以上准确率,显著超越传统单语训练模型。该研究为跨语言身份识别、多语视频理解等应用提供关键技术支持。

原文摘要 · Abstract (English)

Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Challenge, held at ICASSP 2026, focuses on developing methods for face-voice association that are effective when the language at test-time is different than the training one. This report provides a brief summary of the challenge.

跨语言人脸声音多模态身份识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。