arXiv:2604.11110cs.SD2026-04

首个藏语多方言端到端语音大模型,解决数据少、方言差异大难题。

Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan

论文配图:Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
图 1 · 摘自论文原文
  • 用动态Q-Former适配器提取可变长度语音特征,实现跨模态稳定对齐。
  • 在三大藏语方言上达成领先性能,语音识别与语音翻译均最优。
  • 适合低资源语言研究者,为方言协同建模提供可扩展范式。

近年来,语音大语言模型(Speech-LLMs)取得显著进展,极大提升了多模态交互能力。然而,在低资源和方言多样环境中仍面临挑战。藏语数据严重稀缺,且其主要方言(卫藏、安多、康巴)间存在明显发音差异,是典型例证。本文提出Ti-Audio,首个面向藏语的多方言端到端语音大模型。为高效对齐语音与文本,我们引入动态Q-Former适配器,从变长语音中提取关键声学特征,确保在数据有限条件下跨模态对齐稳定。在数据层面,利用相关方言间的相互助益缓解数据稀缺,并采用温度采样策略最大化这种协同效应。实验表明,Ti-Audio在藏语自动语音识别与语音翻译基准上达到当前最优性能。本工作验证了跨方言协作的有效性,为低资源场景下语音大模型开发提供了可扩展范式。

原文摘要 · Abstract (English)

Recent advances in Speech Large Language Models (Speech-LLMs) have made significant progress, greatly enhancing multimodal interaction capabilities.However, their application in low-resource and dialect-diverse environments still faces challenges. The severe scarcity of Tibetan data, coupled with the phonetic differences among its major dialects (Ü-Tsang, Amdo, and Kham), is a prime example of this challenge. This paper proposes Ti-Audio, the first multi-dialectal end-to-end Speech-LLM for Tibetan. To efficiently align speech and text, we introduce a Dynamic Q-Former Adapter that extracts essential acoustic features from variable-length speech, ensuring stable cross-modal alignment even with limited data. At the data level, we leverage mutual assistance among related dialects to alleviate data scarcity and employ a temperature-based sampling strategy to maximize this synergy. Experimental results demonstrate that Ti-Audio achieves state-of-the-art performance on Tibetan benchmarks for automatic speech recognition and speech translation. Our work validates the effectiveness of cross-dialectal cooperation and provides a scalable paradigm for the development of Speech-LLM in low-resource scenarios.

语音大模型藏语多方言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。