arXiv:2410.03930cs.CLcs.SD2024-10被引 6

开源语音识别与说话人分离模型,性能超越现有开源方案。

Reverb: Open-Source ASR and Diarization from Rev

  • 提供完整生产级管道和精简研究模型供实验使用。
  • 在长时语音识别多个领域表现优于现有开源模型。
  • 适合语音技术研究者与开发者快速落地应用。

我们现开放核心语音识别与说话人分离模型的非商业使用权限。发布内容包括面向开发者的完整生产级流水线,以及用于实验的简化研究模型。这些模型旨在推动语音技术领域的研究与创新。今日发布的语音识别模型在多种长时语音识别任务中,性能均超越现有所有开源模型。

原文摘要 · Abstract (English)

Today, we are open-sourcing our core speech recognition and diarization models for non-commercial use. We are releasing both a full production pipeline for developers as well as pared-down research models for experimentation. Rev hopes that these releases will spur research and innovation in the fast-moving domain of voice technology. The speech recognition models released today outperform all existing open source speech recognition models across a variety of long-form speech recognition domains.

语音识别说话人分离开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。