用合成数据和新投影方法,让小模型高效完成多语种语音任务。
NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

- 用仅需ASR数据的SpeechMapper替换旧投影器,提升语音到大模型嵌入效果。
- 构建人工生成的科学演讲数据集fakACL,增强SQA任务训练效果。
- 模型更小巧且性能更强,与去年冠军系统并列第一。
本文介绍NAVER LABS Europe参与IWSLT 2026指令遵循语音处理短赛道的提交方案。在受限设置下,构建可联合执行英文语音到中文、意大利语和德语的自动语音识别(ASR)、语音翻译(ST)及语音问答(SQA)的系统。基于去年短赛道排名第一的成果,我们更新了多阶段训练流程,以SpeechMapper替代原有语音投影器,该方法仅使用ASR数据即可学习语音到大语言模型(LLM)的嵌入映射。同时,我们引入合成数据集fakACL,由大语言模型生成科学演讲内容,再通过SeamlessM4T-large-v2合成语音。结合改进的语音投影机制与领域专用合成数据,本模型在性能上超越去年最佳系统,且模型规模更小,依赖的LLM backbone也更弱。今年结果使系统在整体短赛道排名中并列第一。
原文摘要 · Abstract (English)
In this paper, we describe NAVER LABS Europe's submission to the instruction-following speech processing short track at IWSLT 2026. We participate again in the constrained setting, developing systems capable of jointly performing ASR, ST, and SQA from English speech into Chinese, Italian, and German. Building on our previous submission, ranked first in last year's short track, we update our multi-stage training pipeline by replacing the speech projector with SpeechMapper, a method for learning a speech-to-LLM embedding projector using only ASR data. In addition, we introduce a synthetic SQA dataset, fakACL, composed of artificially generated scientific presentations. This dataset is built by prompting the LLM backbone, segmenting the generated talks, and synthesizing speech with SeamlessM4T-large-v2. The combination of an improved speech projection mechanism and domain-specific synthetic data allows our model to outperform last year's best short-track system, while being considerably more compact and relying on a weaker LLM backbone. This year's results place our system tied for first place in the overall short track ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。