arXiv:2606.27862cs.CV2026-06

解决神经网络多尺度建模中的频率干扰问题,提升图像音频3D重建精度。

ScaLe-INR: Scale and Learn Implicit Neural Representations

论文配图:ScaLe-INR: Scale and Learn Implicit Neural Representations
图 1 · 摘自论文原文
  • 采用多分支结构,按频率分布匹配不同尺度信号的表达区域。
  • 在图像重建上提升5.16 dB,音频重建达50.02 dB,3D重建交并比0.999。
  • 通过边缘引导损失实现分支解耦,防止高频信息干扰低频结构。

隐式神经表示(INRs)由多层感知机参数化,擅长建模连续信号。然而,其核心挑战在于谱偏置与信息串扰:当单一网络试图捕捉多尺度现象时,高频权重更新会破坏底层低频结构近似。我们提出新型多分支架构ScaLe-INR,通过将信号频谱与INR最优工作区显式匹配,解决上述限制。基于傅里叶逆缩放定理,我们证明方向性坐标缩放可扩展网络在特定空间轴上的表征带宽。为数学上强制功能解耦并最小化分支间任务特异性信息泄露,提出方向边缘引导损失,一种源自真实梯度的空间条件稀疏先验。约束高频分支仅作为局部边缘滤波器,消除频谱串扰,加速收敛,并在复杂多尺度拓扑上实现高保真信号重建。我们在多种重建与反演任务中评估ScaLe-INR,性能显著优于现有最先进方法。相比最近基线,图像重建提升+5.16 dB,图像去噪提升+0.65 dB;音频重建达50.02 dB,3D重建交并比0.999,全面超越所有最先进模型。

原文摘要 · Abstract (English)

Implicit Neural Representations (INRs) parameterized by multilayer perceptrons excel at modeling continuous signals. However, a key challenge persists as INRs fundamentally suffer from spectral bias and information cross-talk. When a single network attempts to capture multi-scale phenomena, high-frequency weight updates destructively interfere with the underlying low-frequency structural approximation. We introduce Scale and Learn INR (ScaLe-INR), a novel multi-branch architecture that resolves these limitations by explicitly matching the signal's frequency spectrum with the optimal operating region of the INR. Drawing upon the Fourier inverse scaling theorem we demonstrate that applying directional coordinate scaling expands a network's representational bandwidth along specific spatial axes. To mathematically enforce functional disentanglement and minimize task-specific information leakage between branches, we propose a Directional Edge Guidance Loss, a spatially-conditioned sparsity prior derived from ground-truth gradients. By constraining the high-frequency branches to act as strict, localized edge-filters, ScaLe-INR eliminates spectral cross-talk, accelerates convergence, and achieves high-fidelity signal reconstruction on complex multi-scale topologies. We evaluate ScaLe-INR across diverse reconstruction and inverse tasks, demonstrating substantial performance gains over existing state-of-the-art (SOTA) methods. The proposed architecture improves upon the nearest baselines by +5.16 dB in image reconstruction and +0.65 dB in image denoising. Furthermore, it achieve an impressive figure of 50.02 dB on audio reconstruction and 0.999 IOU(Intersection Over Union) on 3D reconstruction which beats the all SOTA models.

隐式表示多尺度建模图像重建3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。