挑战跨语言跨模态缺失下的说话人识别难题
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge

- 设计真实场景下多语言、缺模态的说话人识别测试基准
- 验证模型在音频/视频缺失时仍能保持识别性能
- 适合研究鲁棒性与泛化能力的多模态系统者
多模态说话人识别系统通常假设训练和测试阶段具备完整且同质的音视频模态,且每位说话人仅使用一种语言。然而在真实应用中,这些假设常不成立:因遮挡、设备故障或隐私限制导致音视频信息缺失;多语言说话人带来语言差异带来的额外复杂性。这些情况严重挑战了系统的鲁棒性与泛化能力。POLY-SIM 2026挑战旨在应对上述问题,提供标准化评估框架,以比较各类解决方案在复杂条件下的表现。
原文摘要 · Abstract (English)
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing, and assume each speaker only speaks a single language. However, in real-world applications, such assumptions often do not hold. Visual or audio information may be missing due to occlusions, camera or microphone failures, or privacy constraints. Multilingual speakers introduce additional complexity due to linguistic variability across languages. These situations constitute substantial challenges for the robustness and generalization capabilities of multimodal speaker identification systems. Aim of the POLY-SIM 2026 challenge is to address these aspects of speaker identification and to provide a standardized setup for the comparison of the proposed solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。