提出因果层选择方法,让音频生成更准
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
- 用前向门控消融法找出驱动生成的关键层
- 在语音和通用音频上性能优于传统对齐方法
- 适合做音频生成与流模型优化的研究者
表示对齐(REPA)通过将生成流模型的中间隐藏状态与预训练教师特征对齐来提升训练效果,但在基于标记条件的音频流模型中,其性能高度依赖于监督层的选择,而这一选择通常基于深度进行启发式判断。本文提出归属引导表示对齐(AG-REPA),一种用于音频流模型中的因果层选择策略。我们发现:在教师空间中语义/声学信息存储能力强的层,并不一定是驱动生成速度场贡献最大的层,这种现象称为‘存储-贡献解耦’(SCD)。为将此洞察转化为可操作的训练指导,我们提出前向仅门控消融(FoG-A),通过观察每层被移除后预测速度场的变化来量化其因果贡献,实现稀疏层选择与自适应加权对齐。在统一语音与通用音频训练(LibriSpeech + AudioSet)及不同标记条件拓扑下,AG-REPA均稳定超越基线。结果表明,对驱动速度场的因果主导层进行对齐,比对表征丰富但功能被动的层对齐更有效。
原文摘要 · Abstract (English)
REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。