复杂神经网络在特定信号任务中更优,非普遍适用。
When do complex-valued neural networks help? A study of representation, geometry, and optimization

- 从表示、几何与优化角度对比复数与实数模型
- 相位任务中复数模型优势明显,幅度任务则实数更优
- 适合信号处理、量子计算等具有相位对称性的场景
复数神经网络(CVNNs)常用于自然以幅值和相位编码信息的领域。但仅含复数输入并不决定学习性能提升:标签信号可能存在于幅值、相位、其耦合或对称性中,而实数模型在合适坐标下也能表达。本文通过代表性的复数、笛卡尔实数、极坐标、纯相位、纯幅值、参数匹配实数及浮点运算匹配实数基线,在合成射频任务上进行评估。结果表明,复数表示有优势但非普适;仅相位调制(PSK)任务偏好相位感知与复数模型,仅正交幅度调制(QAM)任务偏好幅值模型,混合调制时复数优势微弱,未见载波相位旋转则破坏依赖坐标的模型。在量子波函数预测中,动量仅可通过相位恢复;脑电分析信号实验显示,相位锁定、幅值突变、相位-幅值耦合各支持不同坐标视角。发现RadioML 2018.01A存在基准测试偏差:在共享试验选择下,复数CReLU模型比最优实数基线高22.94个百分点;但在每族独立调参、16次试验搜索空间下,差距缩小至2.46个百分点。梯度分析表明,虚高学习率导致实数基线首步不稳定,而复数参数耦合使损失信号分布更稳健。学习率×激活函数因子验证失败主要源于超参数设计。总体而言,CVNNs是依赖于表示、对称性与优化的结构化归纳偏置,而非普遍优越架构。
原文摘要 · Abstract (English)
Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued inputs alone do not determine when complex arithmetic improves learning: the label signal may lie in amplitude, phase, their coupling, or a symmetry that real-valued models can also represent under suitable coordinates. We study this through a representation-first evaluation of CVNNs against Cartesian real, polar, phase-only, magnitude-only, parameter-matched real, and FLOP-matched real baselines. Across synthetic RF tasks, complex representations are useful but not universally superior. PSK-only tasks favor phase-aware and complex-valued models, QAM-only tasks favor magnitude-based models, mixed PSK+QAM gives only a small complex-valued advantage, and unseen carrier-phase rotations break coordinate-dependent models without augmentation. Similar patterns appear beyond RF: in quantum-wavefunction prediction, momentum is invisible to $|ψ|$ but recoverable from phase, while EEG analytic-signal experiments show that phase locking, amplitude bursts, and phase-amplitude coupling each favor different coordinate views. We also identify a benchmarking artifact on RadioML 2018.01A. Under matched-shared-trial selection, a CReLU complex model exceeds the best real baseline by 22.94 PP; under independent per-family tuning on the same data and 16-trial search space, the gap collapses to 2.46 PP. Gradient analysis traces the inflated gap to high-learning-rate first-step instability in real baselines, while complex parameter coupling distributes the loss signal more robustly. A learning-rate $\times$ activation factorial confirms the failure is primarily hyperparameter-driven. Overall, CVNNs are best viewed as structured inductive biases whose gains depend on representation, symmetry, and optimization, not as universally superior architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。