arXiv:2605.10470cs.CV2026-05

提出可证明的多模态超分框架,提升图像重建的语义一致性与泛化能力。

Adaptive Context Matters: Towards Provable Multi-Modality Guidance for Super-Resolution

论文配图:Adaptive Context Matters: Towards Provable Multi-Modality Guidance for Super-Resolution
图 1 · 摘自论文原文
  • 设计动态加权与自适应温度调度,实现时空可变的多模态融合。
  • 理论证明优化模态权重与贡献对齐可降低泛化风险,提升性能。
  • 适用于需高保真语义重建的图像超分辨率任务。

超分辨率(SR)是一个公认的病态问题,存在固有歧义。尽管近期的语义引导和多模态SR方法利用大模型或外部先验增强语义对齐,但异构模态融合在实践与理论上仍不充分。本文首次建立多模态SR的理论模型,揭示先前方法受限于次优的模态利用。分析表明,通过强化模态权重与其有效贡献的对齐,并降低表示复杂度,可改善泛化风险界。这一理论洞察催生了新型多模态专家混合超分辨率框架(M$^3$ESR),采用面向泛化的动态模态融合策略,实现精准风险控制与模态贡献优化。具体包括新颖的空间动态模态加权模块与时间自适应温度调度机制,支持灵活的时空模态加权以实现有效风险控制。大量实验表明,所提M$^3$ESR显著提升泛化性与语义一致性表现,验证其优越性。

原文摘要 · Abstract (English)

Super-resolution (SR) is a severely ill-posed problem with inherent ambiguity, as widely recognized in both empirical and theoretical studies. Although recent semantic-guided and multi-modal SR methods exploit large models or external priors to enhance semantic alignment, the fusion of heterogeneous modalities remains insufficiently understood in practice and theory. In this work, we provide the first theoretical modeling of multi-modal SR, revealing that prior methods are bottlenecked by sub-optimal modality utilization. Our analysis shows that the generalization risk bound can be improved by strengthening the alignment between modality weights and their effective contributions, while reducing representation complexity. This theoretical insight inspires us to propose the novel Multi-Modal Mixture-of-Experts Super-Resolution framework (M$^3$ESR) that employs generalization-oriented dynamic modality fusion for accurate risk control and modality contribution optimization. In detail, we propose a novel spatially dynamic modality weighting module and a temporally adaptive modality temperature scheduling mechanism, enabling flexible and adaptive spatial-temporal modality weighting for effective risk control. Extensive experiments demonstrate that our M$^3$ESR significantly boosts generalization and semantic consistency performances, which confirms our superiority.

超分辨率多模态理论分析动态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。