arXiv:2502.00869cs.CV2025-02被引 5

提出可学习正弦激活函数,提升隐式神经表示的细节建模能力。

A Unified Theory of Sinusoidal Activation Families for Implicit Neural Representations

  • 设计可学习振幅、频率和相位的正弦激活函数STAF。
  • 在图像、音频、三维重建等任务上显著提升PSNR/SSIM指标。
  • 适合需要高精度重建与参数效率的隐式神经表示场景。

隐式神经表示(INRs)以紧凑神经网络建模连续信号,已成为视觉、图形学与信号处理的标准工具。核心挑战在于不依赖复杂手工编码或脆弱训练技巧的前提下精准捕捉细粒度特征。现有研究中,周期性激活函数展现出优越性:从使用单个固定频率正弦的SIREN,到采用多正弦并可训练频率与相位的最新架构。本文系统研究此类正弦激活家族,建立可学习正弦激活的理论与实践框架。具体地,提出傅里叶类激活函数STAF,其振幅、频率和相位均可学习。分析表明:(i) 构建了与标准正弦网络的克罗内克等价结构,量化表达能力增长;(ii) 揭示了可学习正弦参数化下神经正切核(NTK)谱的变化规律;(iii) 提出初始化方法,使激活后输出呈标准正态分布,无需渐近中心极限定理支撑。实验显示,在图像、音频、形状、反问题(超分辨率、去噪)及NeRF任务中,STAF在各类重建指标(如PSNR/SSIM)上表现优异,且层共享下参数效率更高。尽管周期激活缓解了频谱偏见的表现,但未消除其本质;可学习正弦能改善容量-优化权衡,在评估设置中更具优势。

原文摘要 · Abstract (English)

Implicit Neural Representations (INRs) model continuous signals with compact neural networks and have become a standard tool in vision, graphics, and signal processing. A central challenge is accurately capturing fine detail without heavy hand-crafted encodings or brittle training heuristics. Across the literature, periodic activations have emerged as a compelling remedy: from SIREN, which uses a single sinusoid with a fixed global frequency, to more recent architectures employing multiple sinusoids and, in some cases, trainable frequencies and phases. We study this family of sinusoidal activations and develop a principled theoretical and practical framework for trainable sinusoidal activations in INRs. Concretely, we instantiate this framework with Sinusoidal Trainable Activation Functions (STAF), a Fourier-like activation whose amplitudes, frequencies, and phases are learned. Our analysis (i) establishes a Kronecker-equivalence construction that expresses trainable sinusoidal activations with standard sine networks and quantifies expressive growth, (ii) characterizes how the Neural Tangent Kernel (NTK) spectrum changes under trainable sinusoidal parameterization, and (iii) provides an initialization that yields standard normal post-activations without asymptotic central limit theorem (CLT) arguments. Empirically, on images, audio, shapes, inverse problems (super-resolution, denoising) and NeRF, STAF is competitive and often stronger on distortion-oriented reconstruction metrics such as PSNR/SSIM across the evaluated INR tasks, with favorable parameter efficiency under layer-wise sharing. While periodic activations can alleviate practical manifestations of spectral bias, our results indicate they do not eliminate it; instead, trainable sinusoids can improve the observed capacity-optimization trade-off in the evaluated settings.

隐式表示正弦激活神经渲染可学习参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。