arXiv:2601.17723cs.CV2026-01

对比6种隐式表示模型,发现复杂度提升对超分效果帮助有限。

Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark

  • 统一训练条件测试6种隐式神经表示方法
  • 辅助目标可显著提升纹理保真度,优于传统L1损失
  • 模型性能高度依赖训练配置,需系统化评估

隐式神经表示(INR)已成为任意尺度图像超分辨率(ASSR)的标准方法。然而,尚无系统性实证研究在一致条件下评估现有方法的有效性,也未深入探究不同训练策略(如目标设计、优化方式、缩放行为)的影响。我们通过在9种受控训练配方下训练6种基于INR的ASSR方法,并在7个数据集和7种图像质量评估指标上进行评估,构建统一排名框架,实现更可靠的性能比较与结论解读。此外,我们分析了精细控制的训练配置对感知质量的影响,以及辅助目标在保留边缘、纹理和细节中的作用。结果揭示:(1)近期更复杂的INR方法相比早期方法仅带来微小提升,表明当前架构在基准上已趋于饱和;(2)模型性能与训练配置强相关,此前研究常忽略此因素;(3)辅助目标普遍提升纹理保真度,优于标准L1损失,凸显目标设计对特定感知增益的关键作用;(4)INR-based ASSR在模型容量、训练算力和数据多样性上表现出一致且单调的缩放行为,但随复杂度增加收益递减。

原文摘要 · Abstract (English)

Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). However, no systematic empirical study has examined the effectiveness of existing methods under consistent conditions, nor investigated the effects of different training recipes, such as objective design, optimization strategies, and scaling behavior. A rigorous empirical analysis is essential not only for benchmarking performance and revealing true gains but also for establishing the current state of ASSR, identifying saturation limits, and highlighting promising directions. We fill this gap by training 6 INR-based ASSR methods under 9 controlled training recipes and evaluating them across 7 datasets and 7 image quality assessment (IQA) metrics, presenting aggregated performance results through a unified ranking framework that enables more reliable interpretation of performance comparisons and evaluation claims. Furthermore, we investigate the impact of carefully controlled training configurations on perceptual image quality and analyze the role of auxiliary objectives in preserving edges, textures, and fine details during training. Our analysis yields the following insights, previously overlooked: (1) Recent, more complex INR methods provide only marginal improvements over earlier methods, indicating architectural saturation on existing benchmarks. (2) Model performance is strongly correlated with training configurations, a factor neglected in prior comparisons. (3) Auxiliary objectives consistently enhance texture fidelity across architectures compared to standard L1-Loss, emphasizing the role of objective design for targeted perceptual gains. (4) INR-based ASSR exhibits consistent, monotonic scaling behavior across model capacity, training compute, and data diversity, though with diminishing returns as complexity grows.

超分辨率隐式表示训练策略图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。