arXiv:2507.22404cs.CVcs.AI2025-07中稿 · ICCV被引 1

用隐式神经表示结合掩码图像建模,提升图像重建的鲁棒性。

MINR: Implicit Neural Representations with Masked Image Modelling

  • 将隐式神经表示与掩码图像建模结合,学习连续图像函数
  • 在分布内和分布外数据上均优于MAE,且模型更轻量
  • 适合追求高效鲁棒自监督学习的研究者

自监督学习方法如掩码自编码器(MAE)在基于图像重建的预训练任务中表现出显著潜力,但其性能通常高度依赖训练时的掩码策略,且在分布外数据上表现下降。为此,我们提出掩码隐式神经表示(MINR)框架,将隐式神经表示与掩码图像建模相结合。MINR通过学习连续函数表示图像,实现了对掩码策略不敏感的更鲁棒、泛化性更强的重建。实验表明,MINR不仅在分布内场景下优于MAE,也在分布外设置中表现更优,同时降低了模型复杂度。MINR的通用性扩展至多种自监督学习应用,验证了其作为现有框架稳健高效替代方案的实用性。

原文摘要 · Abstract (English)

Self-supervised learning methods like masked autoencoders (MAE) have shown significant promise in learning robust feature representations, particularly in image reconstruction-based pretraining task. However, their performance is often strongly dependent on the masking strategies used during training and can degrade when applied to out-of-distribution data. To address these limitations, we introduce the masked implicit neural representations (MINR) framework that synergizes implicit neural representations with masked image modeling. MINR learns a continuous function to represent images, enabling more robust and generalizable reconstructions irrespective of masking strategies. Our experiments demonstrate that MINR not only outperforms MAE in in-domain scenarios but also in out-of-distribution settings, while reducing model complexity. The versatility of MINR extends to various self-supervised learning applications, confirming its utility as a robust and efficient alternative to existing frameworks.

隐式表示自监督学习图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。