arXiv:2609.07549cs.CL2026-09

Qwen-Audio-3.0-ASR用专家混合架构实现多语言方言语音识别,支持生产级实用功能。

Qwen-Audio-3.0-ASR Technical Report

论文配图:Qwen-Audio-3.0-ASR Technical Report
图 1 · 摘自论文原文
  • 基于Qwen大模型与专家混合(MoE)结构,统一处理语音识别任务。
  • 支持30种语言、16种中文方言,单次转录即可完成校正与上下文建模。
  • 适合需要高精度多语种、方言及热词定制的工业场景应用。

近年来,自动语音识别(ASR)在数据规模、模型规模与大语言模型(LLMs)深度结合三方面取得显著进展。然而,学术基准性能与真实生产需求之间的差距依然存在,尤其在处理多样区域方言、动态实体与热词、长距离上下文信息及非流畅口语方面。本文介绍 Qwen-Audio-3.0-ASR,一种基于大语言模型的专家混合(MoE)ASR系统,通过统一的指令遵循框架应对生产级挑战。该模型以 Qwen 为基础架构,训练于数千万小时的大规模语音数据。支持跨30种语言和16种中国方言(覆盖八大方言区)。除多语言与方言识别外,还具备行业实体识别、分层热词定制、原生单通转录润色及长音频上下文建模等生产级能力。我们进一步开发了低延迟流式版本 Qwen-Audio-3.0-ASR-Streaming,适用于对时延敏感的应用。在中文、英文、多语言及真实工业测试集上的广泛评估表明,其在多种评测条件下达到或接近顶尖水平,性能优于GPT-4o Transcribe与Gemini 3.1 Pro等主流商业系统。

原文摘要 · Abstract (English)

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

语音识别大模型多语言方言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。