用可学习的坐标变换提升隐式神经表示的精度与效率
Scaling Implicit Fields via Hypernetwork-Driven Multiscale Coordinate Transformations
- 通过超网络动态生成多尺度坐标变换,解耦信号复杂结构
- 在图像、3D形状重建上实现4倍更高精度,参数量减少30%-60%
- 适合需要高保真低参数模型的研究者和工业应用
隐式神经表示(INRs)已成为图像、3D形状、符号距离场和辐射场等信号建模的强大范式。尽管在架构设计(如SIREN、FFC、基于KAN的INRs)和优化策略(元学习、摊销、蒸馏)方面取得进展,现有方法仍存在两大核心局限:(1) 表示瓶颈,迫使单一MLP统一建模异质局部结构;(2) 可扩展性受限,缺乏能动态适应信号复杂度的分层机制。本文提出超坐标隐式神经表示(HC-INR),通过超网络学习信号自适应的坐标变换,打破表示瓶颈。HC-INR将表征任务分解为两部分:(i) 学习的多尺度坐标变换模块,将输入域映射至解耦的潜在空间;(ii) 紧凑的隐式场网络,在变换后空间中以显著降低的复杂度建模信号。所提模型采用分层超网络架构,根据局部信号特征条件化坐标变换,实现表示容量的动态分配。理论上证明,HC-INR严格提升可表示频带的上界,同时保持Lipschitz稳定性。大量实验表明,在图像拟合、形状重建和神经辐射场逼近任务中,HC-INR相比强基线实现最高4倍的重建保真度提升,且参数量减少30%–60%。
原文摘要 · Abstract (English)
Implicit Neural Representations (INRs) have emerged as a powerful paradigm for representing signals such as images, 3D shapes, signed distance fields, and radiance fields. While significant progress has been made in architecture design (e.g., SIREN, FFC, KAN-based INRs) and optimization strategies (meta-learning, amortization, distillation), existing approaches still suffer from two core limitations: (1) a representation bottleneck that forces a single MLP to uniformly model heterogeneous local structures, and (2) limited scalability due to the absence of a hierarchical mechanism that dynamically adapts to signal complexity. This work introduces Hyper-Coordinate Implicit Neural Representations (HC-INR), a new class of INRs that break the representational bottleneck by learning signal-adaptive coordinate transformations using a hypernetwork. HC-INR decomposes the representation task into two components: (i) a learned multiscale coordinate transformation module that warps the input domain into a disentangled latent space, and (ii) a compact implicit field network that models the transformed signal with significantly reduced complexity. The proposed model introduces a hierarchical hypernetwork architecture that conditions coordinate transformations on local signal features, enabling dynamic allocation of representation capacity. We theoretically show that HC-INR strictly increases the upper bound of representable frequency bands while maintaining Lipschitz stability. Extensive experiments across image fitting, shape reconstruction, and neural radiance field approximation demonstrate that HC-INR achieves up to 4 times higher reconstruction fidelity than strong INR baselines while using 30--60\% fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。