arXiv:2601.18184cs.SDcs.AI2026-01被引 21

一句话搞定60分钟长音频的语音识别与说话人分离。

VIBEVOICE-ASR Technical Report

  • 端到端统一处理语音识别、说话人分离和时间戳,支持单次输入60分钟音频。
  • 可处理50多种语言及跨语种混用,无需设置语言参数。
  • 通过提示注入自定义上下文,提升专业术语和多音字识别准确率。

本报告介绍 VibeVoice-ASR,一个基于 VibeVoice 的通用语音理解框架,旨在解决尽管短时语音识别技术取得进展,长音频(如会议、播客)中仍存在的上下文碎片化与多说话人复杂性问题。与依赖音频分块的传统流水线方法不同,VibeVoice-ASR 支持长达60分钟音频的单次流式处理,将自动语音识别、说话人分离与时间戳生成统一为一个端到端生成任务。该框架支持超过50种语言,无需显式指定语言,原生支持句内及句间语码转换。此外,我们提出一种基于提示的上下文注入机制,允许用户输入定制化上下文,显著提升领域专有术语与多音字辨识的准确性。

原文摘要 · Abstract (English)

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g., meetings, podcasts) that remain despite recent advancements in short-form speech recognition. Unlike traditional pipelined approaches that rely on audio chunking, VibeVoice-ASRsupports single-pass processing for up to 60 minutes of audio. It unifies Automatic Speech Recognition, Speaker Diarization, and Timestamping into a single end-to-end generation task. In addition, VibeVoice-ASR supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Furthermore, we introduce a prompt-based context injection mechanism that allows users to supply customized conetxt, significantly improving accuracy on domain-specific terminology and polyphonic character disambiguation.

语音识别长音频多说话人端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。