arXiv:2508.04592cs.CV2025-08被引 7

研究多语言环境下人脸与声音的关联,提升跨模态识别能力。

Face-voice Association in Multilingual Environments (FAME) 2026 Challenge Evaluation Plan

  • 构建多语言音视频数据集 MAV-Celeb,模拟真实多语环境。
  • 设计挑战任务评估模型在多语言场景下的跨模态匹配性能。
  • 适合对多模态学习、跨语言识别感兴趣的学者和工程师。

技术进步推动了多模态系统在现实应用中的广泛使用,其中音视频系统尤为普遍。近年来,由于人脸与声音之间存在独特关联,二者匹配问题受到关注。2026年多语言环境下人脸-声音关联(FAME 2026)挑战聚焦于多语言场景下的音视频关联研究。该场景源于全球一半人口为双语者,多数人处于多语言交流情境。挑战使用名为 Multilingual Audio-Visual (MAV-Celeb) 的数据集,探索多语言环境中的脸声关联。本报告详述挑战任务、数据集、基线模型及评估细节。

原文摘要 · Abstract (English)

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, audio-visual systems are among the most widely used multimodal systems. In the recent years, associating face and voice of a person has gained attention due to the presence of unique correlation between them. The Face-voice Association in Multilingual Environments (FAME) 2026 Challenge focuses on exploring face-voice association under the unique condition of a multilingual scenario. This condition is inspired from the fact that half of the world's population is bilingual and most often people communicate under multilingual scenarios. The challenge uses a dataset named Multilingual Audio-Visual (MAV-Celeb) for exploring face-voice association in multilingual environments. This report provides the details of the challenge, dataset, baseline models, and task details for the FAME Challenge.

多模态语音识别跨语言人脸匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。