提出离散流匹配方法,实现高效并行语音识别
Drax: Speech Recognition with Discrete Flow Matching
- 用音频条件概率路径引导模型学习真实推理错误轨迹
- 在相同精度下比现有模型更快,提升识别效率
- 适合追求高吞吐量语音识别的工程应用
扩散和基于流的非自回归(NAR)模型在大语言建模中表现优异,但在自动语音识别(ASR)中的潜力尚未被充分探索。我们提出Drax,一种用于ASR的离散流匹配框架,支持高效并行解码。为更好对齐训练与推理,我们构建了音频条件概率路径,引导模型沿类似可能中间推理错误的轨迹演进,而非直接从随机噪声到目标转换。理论分析表明,泛化差距与训练和推理状态分布差异相关,受累积速度误差控制,从而支持我们的设计选择。实证评估显示,该方法在识别准确率上达到当前最优水平,同时提供更优的精度-效率权衡,凸显离散流匹配在推进NAR ASR方面的前景。
原文摘要 · Abstract (English)
Diffusion and flow-based non-autoregressive (NAR) models have shown strong promise in large language modeling, however, their potential for automatic speech recognition (ASR) remains largely unexplored. We propose Drax, a discrete flow matching framework for ASR that enables efficient parallel decoding. To better align training with inference, we construct an audio-conditioned probability path that guides the model through trajectories resembling likely intermediate inference errors, rather than direct random noise to target transitions. Our theoretical analysis links the generalization gap to divergences between training and inference occupancies, controlled by cumulative velocity errors, thereby motivating our design choice. Empirical evaluation demonstrates that our approach attains recognition accuracy on par with state-of-the-art speech models while offering improved accuracy-efficiency trade-offs, highlighting discrete flow matching as a promising direction for advancing NAR ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。