arXiv:2604.08384eess.AScs.AI2026-04

通过可控的语音对齐模拟,提升低资源语音大模型训练效果

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

  • 根据指定错误率范围模拟CTC后验分布,实现对语音对齐难度的精准控制
  • 在多种迁移场景下显著提升语音识别准确率,优于TASU和基于TTS的方法
  • 无需语音合成即可优化训练过程,适合低资源语音模型微调

语音大模型的后训练越来越依赖高效的跨模态对齐与鲁棒的低资源适应,但大规模音文配对数据的收集仍成本高昂。仅用文本对齐的方法如TASU通过从文本中模拟CTC后验来降低负担,但其对不确定性和错误率的控制有限,课程设计多为经验性。我们提出TASU2,一个可控制的CTC模拟框架,在指定的词错误率(WER)范围内生成CTC后验分布,从而产生更贴合声学解码接口的文本衍生监督信号。这使得无需语音合成即可构建系统性的后训练课程,平滑调节监督难度。在多个源到目标的适应设置中,TASU2在域内与域外识别性能上均优于TASU,且持续超越强基线方法,包括纯文本微调和基于TTS的数据增强,同时缓解了源域性能退化问题。

原文摘要 · Abstract (English)

Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU reduce this burden by simulating CTC posteriors from transcripts, but they provide limited control over uncertainty and error rate, making curriculum design largely heuristic. We propose \textbf{TASU2}, a controllable CTC simulation framework that simulates CTC posterior distributions under a specified WER range, producing text-derived supervision that better matches the acoustic decoding interface. This enables principled post-training curricula that smoothly vary supervision difficulty without TTS. Across multiple source-to-target adaptation settings, TASU2 improves in-domain and out-of-domain recognition over TASU, and consistently outperforms strong baselines including text-only fine-tuning and TTS-based augmentation, while mitigating source-domain performance degradation.

语音大模型对齐训练低资源学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。