FBK提出长/短语音指令模型,长语音30秒分段最稳定,效果领先。
FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

- 长语音分段采用固定30秒策略,提升生成稳定性。
- 短语音任务SIFS达2.0708,长语音HIFS最高2.0663。
- 模型扩展后仍保持短语音能力,重复插入是主要幻觉形式。
本文介绍我们在IWSLT 2026指令遵循共享任务中的提交方案。针对资源受限场景,我们开发了短语音与长语音指令遵循的SpeechLLMs。在短语音赛道,模型在MCIF数据集上取得2.0708的SIFS得分。在长语音赛道,探索了三种语音分割方法,并引入HIFS评分以应对长序列生成不稳问题。实验表明,固定30秒分割策略表现最优,获得最高HIFS分数2.0663。进一步分析发现,幻觉主要表现为输出中重复插入内容,显著影响ASR和SSUM指标,而长语音扩展并未损害短语音能力。
原文摘要 · Abstract (English)
This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs are developed for both short-form and long-form speech instruction following under constrained settings. For the short track, strong performance is achieved on MCIF, with a SIFS score of 2.0708. For the long track, three speech segmentation methods are explored, and the HIFS score is introduced to account for unstable long-form generation. Experimental results show that fixed 30-second segmentation provides the most robust long-form performance, achieving the highest HIFS score of 2.0663. Further analysis shows that hallucination mainly manifests as repetitive insertions in generated outputs, substantially affecting ASR and SSUM, while short-form capabilities are largely retained after long-form extension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。