arXiv:2601.21337cs.CLcs.SD2026-01被引 99

Qwen3-ASR推出多语言语音识别与对齐模型,性能超越开源同行。

Qwen3-ASR Technical Report

论文配图:Qwen3-ASR Technical Report
图 1 · 摘自论文原文
  • 基于大模型架构的全语言端到端语音识别与非自回归对齐方法
  • 1.7B模型达开源最优,0.6B模型实现92ms首字延迟与每秒处理2000秒语音
  • 支持52种语言,开源发布,适合语音研究与工业部署

本文介绍Qwen3-ASR系列,包含两个强大的全功能语音识别模型和一个新型非自回归语音强制对齐模型。Qwen3-ASR-1.7B和Qwen3-ASR-0.6B支持52种语言和方言的语音识别与语言识别,依托大规模语音训练数据及基础模型Qwen3-Omni的强大音频理解能力。除公开基准外,我们还进行了全面的内部评估,发现1.7B版本在开源模型中达到最先进水平,性能可媲美最强专有API;0.6B版本则在准确率与效率间取得最佳平衡,平均首词响应时间(TTFT)低至92ms,128并发下可实现每秒转录2000秒语音。Qwen3-ForcedAligner-0.6B是基于大语言模型的非自回归时间戳预测器,可在11种语言上对齐文本-语音对,实验表明其在时间戳精度上优于三大最强对齐模型,且在效率与泛化性方面更具优势。为推动语音识别与音频理解社区发展,所有模型均以Apache 2.0许可证开源。

原文摘要 · Abstract (English)

In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR models that support language identification and ASR for 52 languages and dialects. Both of them leverage large-scale speech training data and the strong audio understanding ability of their foundation model Qwen3-Omni. We conduct comprehensive internal evaluation besides the open-sourced benchmarks as ASR models might differ little on open-sourced benchmark scores but exhibit significant quality differences in real-world scenarios. The experiments reveal that the 1.7B version achieves SOTA performance among open-sourced ASR models and is competitive with the strongest proprietary APIs while the 0.6B version offers the best accuracy-efficiency trade-off. Qwen3-ASR-0.6B can achieve an average TTFT as low as 92ms and transcribe 2000 seconds speech in 1 second at a concurrency of 128. Qwen3-ForcedAligner-0.6B is an LLM based NAR timestamp predictor that is able to align text-speech pairs in 11 languages. Timestamp accuracy experiments show that the proposed model outperforms the three strongest force alignment models and takes more advantages in efficiency and versatility. To further accelerate the community research of ASR and audio understanding, we release these models under the Apache 2.0 license.

语音识别大模型非自回归开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。