用小模型和高质量数据实现语音到文本的统一建模。
Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
- 用两阶段方法对齐语音与文本模态,再微调指令理解能力。
- 在<20亿参数的小型语言模型上实现语音识别与问答效果提升。
- 仅使用CC-BY授权数据+合成数据,适合资源受限场景。
本文介绍了IT-IST团队在IWSLT 2025指令跟随语音处理共享任务中的提交方案。我们参与了短赛道(语音识别、翻译与语音问答)的评测。提出一种统一的语音到文本模型,通过两个阶段实现:首先进行模态对齐,再进行指令微调。关键在于采用小于20亿参数的轻量级语言模型,并仅使用高质量、CC-BY许可的数据,辅以合成数据增强现有资源。该方法在保持模型规模可控的同时,有效提升了多任务表现。
原文摘要 · Abstract (English)
This paper presents the IT-IST submission to the IWSLT 2025 Shared Task on Instruction Following Speech Processing. We submit results for the Short Track, i.e., speech recognition, translation, and spoken question answering. Our model is a unified speech-to-text model that integrates a pre-trained continuous speech encoder and text decoder through a first phase of modality alignment and a second phase of instruction fine-tuning. Crucially, we focus on using small-scale language model backbones (< 2B) and restrict to high-quality, CC-BY data along with synthetic data generation to supplement existing resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。