提出新型复数玻尔兹曼机,显式建模音频信号的幅度与相位关系。
PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates

- 在极坐标下建模复数变量,让相位依赖于幅度
- 在音频数据上优于传统模型,包括深度神经网络
- 适合处理具有物理意义的复数信号,如通信与量子系统
尽管大量数据(如音频信号频谱)天然以复数形式存在,传统机器学习方法常将其简化为实值变量处理。这种简化虽提升计算效率,却丢失了幅度与相位间的内在关联信息。本文提出一种新型玻尔兹曼机(PolarBM),可自然处理极坐标下的复数变量(即幅度-相位表示)。PolarBM定义了复数变量的概率密度函数,其中相位显式依赖于幅度,从而捕捉复数信号中重要的物理关系。为进一步契合人类听觉感知,提出LogPolarBM,将幅度置于对数尺度建模。该扩展导出一种灵活的条件概率密度函数——幂权非中心复高斯分布(PW-NCCG),其边缘幅度分布涵盖Rice、Nakagami和非中心卡方分布作为特例。针对实际应用,还引入受限变体:PolarRBM与LogPolarRBM。实验表明,通过显式建模幅度与相位的依赖关系,所提RBMs在建模精度上优于传统模型,包括深度神经网络。虽然实验聚焦音频信号,但该模型适用于涉及复数数据的广泛科学与工程领域,如无线通信与量子力学。
原文摘要 · Abstract (English)
Although vast amounts of data, such as audio signal spectra, are naturally represented using complex numbers, conventional machine learning methods often simplify complex-domain problems by employing frameworks designed for real-valued variables. While this simplification offers computational benefits, it discards structural information regarding the inherent relationship between amplitude and phase. In this paper, we propose a novel Boltzmann machine (BM), named PolarBM, capable of naturally handling complex-valued variables in the polar coordinate (i.e., an amplitude-phase representation). PolarBM defines a probability density function for complex variables in which the phase explicitly depends on the amplitude, thereby capturing the physically important relationships of complex-valued signals. Furthermore, to process audio signals in accordance with human auditory perception, we propose LogPolarBM, which models amplitude on a logarithmic scale. This extension yields a flexible conditional probability density function, a power-weighted noncentral complex Gaussian (PW-NCCG) distribution, whose marginal amplitude distribution encompasses the Rice, Nakagami, and noncentral chi distributions as special cases. For practical applications, we also introduce the restricted variants of these proposed models: PolarRBM and LogPolarRBM. Experimental results demonstrate that by explicitly modeling the dependency between amplitude and phase, the proposed RBMs achieve superior modeling accuracy compared to conventional models, including deep neural networks. Although our experiments focus on audio signals, the utility of the proposed BMs is not limited to audio applications; their potential extends widely across various fields of science and engineering that involve complex-valued data, such as wireless communications and quantum mechanics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。