用神经隐式场提升多图超分辨率,无需高分辨率训练数据。
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution

- 以坐标神经网络表示图像,共享低分辨率多帧的隐式场
- 联合优化帧对齐与隐式场,支持最高8倍上采样
- 不依赖训练数据,适合卫星与手机拍摄图像增强
高分辨率成像常受限于传感器技术、大气条件和成本,如卫星遥感和手持相机(如手机)。单图超分辨率需强先验或辅助数据,易产生虚假结构。多图超分辨率(MISR)通过子像素偏移的多视角约束提升光学分辨率。本文提出SuperF,一种基于神经隐式场(INR)的测试时优化方法,共享多个低分辨率帧的隐式场,并联合优化帧对齐与场表示。相比现有基线,直接参数化子像素对齐为可优化仿射变换,采用对应输出分辨率的超采样坐标网格进行优化。在模拟卫星影像和地面手持图像上实验,最高支持8倍上采样,效果显著。核心优势:无需任何高分辨率训练数据。
原文摘要 · Abstract (English)
High-resolution imagery is often hindered by limitations in sensor technology, atmospheric conditions, and costs. Such challenges occur in satellite remote sensing, but also with handheld cameras, such as our smartphones. Hence, super-resolution aims to enhance the image resolution algorithmically. Since single-image super-resolution requires solving an inverse problem, such methods must exploit strong priors, e.g. learned from high-resolution training data, or be constrained by auxiliary data, e.g. by a high-resolution guide from another modality. While qualitatively pleasing, such approaches often lead to "hallucinated" structures that do not match reality. In contrast, multi-image super-resolution (MISR) aims to improve the (optical) resolution by constraining the super-resolution process with multiple views taken with sub-pixel shifts. Here, we propose SuperF, a test-time optimization approach for MISR that leverages coordinate-based neural networks, also called neural fields. Their ability to represent continuous signals with an implicit neural representation (INR) makes them an ideal fit for the MISR task. The key characteristic of our approach is to share an INR for multiple shifted low-resolution frames and to jointly optimize the frame alignment with the INR. Our approach advances related INR baselines, adopted from burst fusion for layer separation, by directly parameterizing the sub-pixel alignment as optimizable affine transformation parameters and by optimizing via a super-sampled coordinate grid that corresponds to the output resolution. Our experiments yield compelling results on simulated bursts of satellite imagery and ground-level images from handheld cameras, with upsampling factors of up to 8. A key advantage of SuperF is that this approach does not rely on any high-resolution training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。