提出统一框架,系统梳理语音生成中的自回归解码策略。
A Generalized Formalism of Auto-Regressive Decoding for Speech Processing

- 构建通用理论框架,明确定义自回归解码的范畴与标准。
- 可简化解码策略的基准设计与消融实验,提升可比性。
- 适合研究语音生成、模型推理优化的学者参考。
在语音处理中,大多数先进序列预测模型依赖自回归(AR)策略,基于模型的原始输出逐词生成序列。尽管自回归策略在推理过程中至关重要,但由于其定义隐晦且存在多种解释,目前缺乏对这一领域的系统性综述。这导致策略选择、比较和评估困难,并造成方法分类的不一致。本文首先设定自回归搜索在语音处理中的明确包含标准,进而推导出一个通用理论框架,用于分类和报告神经模型的搜索策略。该形式化方法展示了其在简化以解码过程为中心的基准设计方面的潜力,支持聚焦于搜索策略的消融研究。
原文摘要 · Abstract (English)
In speech processing, most state-of-the-art sequence prediction models rely on auto-regressive (AR) strategies to generate output sequences based on the raw predictions of the model. Despite their crucial role in the inference process, a comprehensive overview of AR strategies as a unified field is lacking, due largely to implicit and multiple definitions of next-token decoding. This context complicates the choice, comparison, and evaluation of strategies, while creating inconsistencies in the characterization of approaches as auto-regressive or not. We begin by setting explicit inclusion criteria for the field of AR search in speech processing, and derive a generalized theoretical framework to categorize and report on search strategies for neural models. We show the capabilities of this formalism in simplifying the design of benchmarks centered around the decoding process, allowing for ablation studies that are focused on search strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。