arXiv:2603.24569cs.CV2026-03被引 3

解决多模态语音识别中缺失视觉信息和跨语言问题的挑战赛设计

POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan

  • 构建在模态缺失与跨语言场景下的评测框架
  • 提供标准化数据集与评估协议,支持不完整输入下的性能测试
  • 适合研究鲁棒性多模态系统或实际应用落地的团队参与

多模态说话人识别系统通常假设训练和测试阶段音频-视觉模态均完整且一致。但在真实场景中,视觉信息可能因遮挡、摄像头故障或隐私限制而缺失,多语言说话人又带来语言差异带来的额外复杂性,严重影响系统的鲁棒性和泛化能力。为此,POLY-SIM Grand Challenge 2026旨在推动在模态缺失与跨语言条件下多模态说话人识别的研究。该挑战鼓励开发能有效利用不完整多模态输入且在不同语言间保持高性能的鲁棒方法。本文介绍了挑战赛的设计与组织,包括数据集、任务设定、评估协议及基线模型。通过提供标准化基准与评估框架,旨在促进更鲁棒、更实用的多模态说话人识别系统的发展。

原文摘要 · Abstract (English)

Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions often do not hold. Visual information may be missing due to occlusions, camera failures, or privacy constraints, while multilingual speakers introduce additional complexity due to linguistic variability across languages. These challenges significantly affect the robustness and generalization of multimodal speaker identification systems. The POLY-SIM Grand Challenge 2026 aims to advance research in multimodal speaker identification under missing-modality and cross-lingual conditions. Specifically, the Grand Challenge encourages the development of robust methods that can effectively leverage incomplete multimodal inputs while maintaining strong performance across different languages. This report presents the design and organization of the POLY-SIM Grand Challenge 2026, including the dataset, task formulation, evaluation protocol, and baseline model. By providing a standardized benchmark and evaluation framework, the challenge aims to foster progress toward more robust and practical multimodal speaker identification systems.

多模态说话人识别跨语言挑战赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。