arXiv:2602.18822cs.CV2026-02

解决真实世界跨模态图像模糊对齐难题,实现高效高质超分辨率重建。

Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations

  • 自监督联合优化对齐翻译器与内容感知参考滤波器
  • 在真实错位数据上达到顶尖性能,速度提升15.3倍
  • 适合需要无标注训练的工业级图像增强场景

真实世界中的跨模态超分辨率面临挑战,因仅有未标注的低分辨率源图像与高分辨率引导图像,且存在复杂的空间错位。以往方法依赖模拟数据或次优对齐策略,忽略跨模态关联,限制实际表现。本文提出RobSelf,一种自监督模型,通过在线联合优化一个对错位敏感的特征翻译器与一个内容感知的参考滤波器。翻译器利用弱监督、错位感知的翻译机制,实现跨模态与跨分辨率对齐,生成对齐的引导特征;在此指导下,滤波器对源图像进行基于参考的判别式自增强,实现高分辨率、高保真度的超分辨率预测。在合成数据与实采集数据上的实验表明,RobSelf性能达到当前最优,优于现有自监督与监督方法。此外,其效率显著提升,相较先前自监督方法最快达15.3倍加速。

原文摘要 · Abstract (English)

Cross-modal super-resolution (SR) on real-world misaligned data is challenging, as only unlabeled low-resolution (LR) source and high-resolution (HR) guide images with complex spatial misalignment are available. Previous methods either rely on simulated training data or adopt suboptimal alignment strategies that overlook cross-modal dependencies, limiting their practical performance. To address these issues, we propose RobSelf, a self-supervised model that jointly optimizes a misalignment-aware feature translator and a content-aware reference filter online. The translator resolves unsupervised cross-modal and cross-resolution alignment via weakly-supervised, misalignment-aware translation, yielding an aligned guide feature. Guided by this feature, the filter performs reference-based discriminative self-enhancement on the source, enabling SR prediction with high resolution and high fidelity. Experiments on synthesized data and collected real-world data demonstrate that RobSelf achieves state-of-the-art performance, outperforming existing self-supervised and supervised methods. Moreover, it achieves superior efficiency, being up to 15.3$\times$ faster than prior self-supervised methods.

超分辨率跨模态自监督图像对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。