arXiv:2606.29335cs.LGcs.AI2026-06被引 1

动态路由多模态信息,解决跨语言语音识别难题。

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

论文配图:AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification
图 1 · 摘自论文原文
  • 根据输入质量动态分配模态权重,实现自适应融合。
  • 跨语言场景下最高准确率达100%,平均99.07%。
  • 适合多语言、缺失模态的实际应用环境。

多模态说话人识别系统在实际部署中面临模态缺失和语言不匹配两大挑战。现实场景中,多说话人背景对话、环境噪声和语音重叠进一步降低识别精度。为此,我们针对POLY-SIM 2026挑战赛提出一种多模态多语种说话人识别系统,核心为自适应模态路由(AMR)模块,可动态评估每样本输入质量并融合模态信息。AMR采用两个模态适配器,分别处理基于语言鲁棒音频编码器(W2V-BERT 2.0)和大规模预训练人脸编码器(IResNet-18)提取的嵌入,生成模态适配嵌入;再通过可训练路由机制估计动态模态权重,并用于聚合模态特定分类结果。为优化路由,采用模态感知训练策略,构建四类样本对模拟多样输入条件,以KL散度作为权重分配的显式监督。在POLY-SIM 2026评测集上,系统在英语多模态(P3)、乌尔都语多模态(P5)、英语单音频(P4)、乌尔都语单音频(P6)四种协议下准确率分别为99.93%、100.00%、97.50%、98.83%,平均准确率达99.07%,较融合与正交投影(FOP)基线提升32.73%。

原文摘要 · Abstract (English)

Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions. In practical scenarios, background multi-speaker conversations, ambient noise, and overlapping speech further degrade identification accuracy. To address these challenges, we propose a multimodal polyglot speaker identification system for the POLY-SIM 2026 Grand Challenge. The system is fundamentally built upon Adaptive Modality Routing(AMR), a modality fusion module that dynamically assesses per-sample input quality and integrates modality information. Specifically, AMR employs two modality adapters to process the embeddings extracted from a linguistically robust audio encoder(W2V-BERT 2.0) and a large-scale pretrained face encoder(IResNet-18), producing modality-adapted embeddings. Based on these adapted embeddings, a trainable router estimates dynamic modality weights, which are subsequently applied to aggregate the modality-specific logits for the final prediction. To optimize this routing mechanism, we adopt a modality-aware training strategy that constructs four types of sample pairs to simulate diverse input conditions, with KL divergence serving as explicit supervision for weight assignment. Experimental results on the POLY-SIM 2026 evaluation set show that the proposed system achieves identification accuracy of 99.93%(English multimodal, P3), 100.00%(Urdu multimodal, P5), 97.50%(English audio-only, P4), and 98.83%(Urdu audio-only, P6). The average accuracy across all four protocols is 99.07%, surpassing the Fusion and Orthogonal Projection(FOP) baseline by 32.73%.

多模态说话人识别自适应路由多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。