用物理约束注意力网络统一生成高质量时频表示,提升语音音乐分析精度。
PHAST-Net: Attention-Guided, Physics-Informed Network for Unified Estimation of Ideal Time-Frequency Representations

- 基于连续对数频率小波族,通过注意力机制融合多视角时频信息。
- 在合成数据上训练,实现高分辨率、抗交叉项干扰的统一时频表示。
- 适合需要精确时频分析的语音、音乐及非平稳信号处理场景。
我们提出 PHAST-Net,一种注意力引导、物理信息嵌入的网络,用于统一估计理想时频表示(ITFRs),涵盖谱图、节奏图、节拍图和和声表示等。该模型从一组连续对数频率自适应小波变换(CLAWT)中学习通用映射,生成高分辨率且交叉项抑制良好的时频表示。所选的 CLAWT 集合基于 Cohen's class 核分析,以最大化对数频率时频平面上的曲率覆盖,契合谐波信号结构。训练中引入物理信息辅助重投影损失,利用预测的 ITFR 和对应的 Cohen's class 核重建理想观测的 CLAWT 集合,促进变换一致性与能量守恒,缓解目标稀疏性问题并增强优化稳定性。注意力层有效抑制跨项干扰。对数频率形式还支持谐波版 PHAST-Net,可分离基频结构,生成鲁棒的仅基频表示,如导出的基频节奏图和节拍图。此外,我们提出 Spline-PHAST-Net,将检测到的时频脊线参数化为连续样条轨迹,支持任意网格重渲染与信号重构。在有效无限量的程序生成数据集上训练,PHAST-Net 在多项指标上优于现有方法,为语音、音乐及更广泛的非平稳信号提供统一的高分辨率、抗交叉项分析框架。
原文摘要 · Abstract (English)
We introduce PHAST-Net, an attention-guided, physics-informed network for unified estimation of Ideal Time-Frequency Representations (ITFRs), spanning spectral, tempo-based, metrical, and harmonic representations such as Spectrograms, Tempograms, and Metrograms. PHAST-Net learns an application-general mapping from a constellation of wavelet transforms, the proposed Continuous Log-frequency Adaptive Wavelet Transform (CLAWT), to high-resolution, cross-term-suppressed time-frequency (T-F) representations. The proposed constellation of CLAWTs is selected through Cohen's class kernel analysis to maximise curvature coverage in a logarithmic-frequency T-F plane tailored to harmonic signal structure. PHAST-Net further incorporates a proposed physics-informed auxiliary reprojection loss designed to reconstruct the idealised observed CLAWT constellation from the predicted ITFR and the corresponding Cohen's class kernels during training. This auxiliary objective promotes transform consistency and energy conservation, mitigates pathological target sparsity, and enhances optimisation stability. Attention layers further promote effective cross-term suppression across the input constellation. The log-frequency formulation also enables Harmonic PHAST-Net, which estimates a Harmonic ITFR that isolates fundamental structure, supporting robust fundamental-only representations for speech and music, such as derived fundamental Tempograms and Metrograms. We further introduce Spline-PHAST-Net, which parameterises detected and associated T-F ridges as continuous spline trajectories, enabling arbitrary-grid re-rendering and signal reconstruction. Trained on an effectively unbounded procedurally generated dataset, PHAST-Net demonstrates improved accuracy over established approaches, providing a unified framework for high-resolution, cross-term-robust analysis of speech, music, and broader nonstationary signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。