arXiv:2503.20212cs.CLeess.AS2025-03被引 11

Dolphin是支持40种东方语言的超大规模语音识别模型,准确率显著超越现有开源方案。

Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages

  • 基于Whisper架构扩展,融合自研与开源数据优化多语言识别
  • 在40种东方语言和22种汉语方言上实现领先准确率
  • 开源模型与代码,助力学术与产业界复现创新

本报告介绍Dolphin,一个大规模多语言自动语音识别(ASR)模型,通过扩展Whisper架构以支持更广泛的语言。该方法整合了内部专有数据与开源数据集,用于优化Dolphin性能。模型专门设计用于在东亚、南亚、东南亚及中东地区的40种东方语言中实现显著的识别准确率,同时支持22种汉语方言。实验评估显示,Dolphin在多种语言上的表现显著优于当前最先进的开源模型。为促进可复现性与社区驱动的创新,我们已公开训练好的模型及推理源代码。

原文摘要 · Abstract (English)

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source datasets to refine and optimize Dolphin's performance. The model is specifically designed to achieve notable recognition accuracy for 40 Eastern languages across East Asia, South Asia, Southeast Asia, and the Middle East, while also supporting 22 Chinese dialects. Experimental evaluations show that Dolphin significantly outperforms current state-of-the-art open-source models across various languages. To promote reproducibility and community-driven innovation, we are making our trained models and inference source code publicly available.

语音识别多语言Whisper开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。