让模型生成多通道相位,而非逐通道恢复,提升音频与地震信号的物理一致性。
RIPPLE: Generating Multi-Channel Phase, Not Recovering It

- 将相位恢复转为生成,用先验引导跨通道相位结构保持
- 地震数据中相位误差从57.3°降至33.8°,显著提升物理合理性
- 适用于空间音频、地震信号等依赖通道间相位关系的任务
生成模型能高保真合成幅度谱,而相位通常由独立于各通道的恢复模块(如Griffin-Lim、声码器或潜在解码器)处理。对于多通道波形,这种分离策略代价高昂:空间音频和三元地震记录中的物理信息存在于通道间的相位关系,而通道独立恢复无法生成此类结构。该损失也难以察觉,因为基于幅度的评价指标在通道间相位相干性崩溃时变化极小——导致流水线丢弃物理信息却仍得高分。本文主张相位应被生成而非恢复,提出RIPPLE(基于先验的跨通道相位修正学习),将Griffin-Lim重新理解为相位先验而非最终估计器:以源相位初始化,携带跨通道结构,通过显式跨通道相位损失下的修正流进行优化。在第一阶全向声场转换与地震站间波形迁移两个物理无关任务上测试,RIPPLE在下游分析所用的相干性指标上均优于传统恢复流水线。地震案例尤为关键:在多种生成架构下,逐通道恢复使S波偏振误差接近57.3°的随机水平,而学习到的相位将其降低至33.8°。
原文摘要 · Abstract (English)
Generative models synthesize magnitude spectra with high fidelity, while phase is delegated to a recovery module---Griffin--Lim, a vocoder, or a latent decoder---applied independently to each channel. For multi-channel waveforms this delegation is costly: the physical content of spatial audio and three-component seismograms lives in the phase relationships between channels, precisely what channel-independent recovery cannot produce. The cost is also invisible, since the magnitude-based metrics common to both fields barely move when inter-channel phase coherence collapses---so a pipeline can discard the physical information in its output while still scoring well. We argue that phase should be generated, not recovered, and present RIPPLE (Rectified Inter-channel Phase with Prior-based LEarning), which reinterprets Griffin--Lim as a phase **prior** rather than a final estimator: initialized from the source phase, this prior carries the inter-channel structure to be preserved, and a rectified flow refines it toward the target under an explicit inter-channel phase loss. Tested on first-order ambisonics environment transfer and seismic cross-station translation---two physically unrelated domains---RIPPLE outperforms recovery-based pipelines on the coherence metrics that downstream analyses consume. The seismic case is decisive: across architecturally distinct generators, per-channel recovery leaves S-wave polarization error near the $57.3^\circ$ random expectation, whereas learned phase reduces it to $33.8^\circ$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。