arXiv:2605.08961cs.CLeess.AS2026-05被引 1

针对中文方言识别难题,打造更小更快的实时语音识别模型

Dolphin-CN-Dialect: Where Chinese Dialects Matter

  • 用温度采样平衡普通话与低资源方言数据
  • 方言识别准确率提升,字符错误率显著降低
  • 模型小巧支持实时推理,适合部署在边缘设备

我们提出Dolphin-CN-Dialect,一个面向中文及方言丰富场景的流式语音识别模型。相比前序版本,该模型在数据处理、分词、训练稳定性与数据采样策略上均有显著改进。为解决方言数据高度不平衡问题,提出基于温度的采样策略,有效平衡标准普通话与低资源方言,显著提升方言识别性能。重新设计分词器,采用汉字级建模(中文)与子词级建模(英文),并引入可扩展的方言标记。实验表明,Dolphin-CN-Dialect在方言识别准确率和字符错误率(CER)方面均优于Dolphin,且性能媲美最新开源SOTA模型,同时模型规模显著更小。支持流式与非流式推理,兼顾延迟与精度。通过热词支持与专用硬件优化,实现灵活定制与高效部署。该模型为实际多方言语音识别应用提供了强大且实用的解决方案。

原文摘要 · Abstract (English)

We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces substantial improvements in data processing, tokenization, training stability, and data sampling strategies. To address the challenges of highly imbalanced dialect data, we propose a temperature-based sampling strategy that effectively balances standard Mandarin and low-resource dialects, leading to significant gains in dialect recognition performance. In addition, we redesign the tokenizer to better align with linguistic characteristics, adopting character-level modeling for Chinese and subword modeling for English, while introducing extensible dialect tokens. Experimental results show that Dolphin-CN-Dialect achieves improvement in dialect recognition accuracy and CER reduction compared to Dolphin. Furthermore, Dolphin-CN-Dialect reaches competitive performance with recent SOTA open-source ASR models, while maintaining a significantly smaller model size. Dolphin-CN-Dialect supports both streaming and non-streaming inference, enabling a practical balance between latency and accuracy. It also provides flexible customization through hotword support and efficient deployment optimized for specialized hardware. These improvements make Dolphin-CN-Dialect a strong and practical solution for real-world multi-dialect ASR applications.

语音识别中文方言流式模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。