arXiv:2504.05795cs.CV2025-04AAAI被引 1

用语言指令精准控制图像融合,应对复杂退化场景

Robust Fusion Controller: Degradation-aware Image Fusion with Fine-grained Language Instructions

  • 通过解析语言指令生成退化类型与空间范围的双重控制条件
  • 在多种复合退化下表现稳健,尤其在强光眩光场景中提升显著
  • 适合需要灵活适应真实环境退化的图像融合应用

现有图像融合方法难以适应包含多样、空间异质性退化的现实环境。为此,我们提出鲁棒融合控制器(RFC),通过细粒度语言指令实现退化感知融合,确保在恶劣环境中的可靠应用。RFC首先解析语言指令,创新性地提取功能条件(需去除的退化类型)和空间条件(退化覆盖范围);随后通过多条件耦合网络生成复合控制先验,实现从抽象语言到潜在控制变量的无缝转换;进而设计基于混合注意力的融合网络,将所得复合控制先验深度嵌入,线性调制中间融合特征;为确保语言指令与控制结果对齐,引入新颖的语言-特征对齐损失,约束特征级增益与复合控制先验的一致性。在公开数据集上的大量实验表明,本方法对多种复合退化具有鲁棒性,尤其在高挑战性的耀斑场景中表现突出。

原文摘要 · Abstract (English)

Current image fusion methods struggle to adapt to real-world environments encompassing diverse degradations with spatially varying characteristics. To address this challenge, we propose a robust fusion controller (RFC) capable of achieving degradation-aware image fusion through fine-grained language instructions, ensuring its reliable application in adverse environments. Specifically, RFC first parses language instructions to innovatively derive the functional condition and the spatial condition, where the former specifies the degradation type to remove, while the latter defines its spatial coverage. Then, a composite control priori is generated through a multi-condition coupling network, achieving a seamless transition from abstract language instructions to latent control variables. Subsequently, we design a hybrid attention-based fusion network to aggregate multi-modal information, in which the obtained composite control priori is deeply embedded to linearly modulate the intermediate fused features. To ensure the alignment between language instructions and control outcomes, we introduce a novel language-feature alignment loss, which constrains the consistency between feature-level gains and the composite control priori. Extensive experiments on publicly available datasets demonstrate that our RFC is robust against various composite degradations, particularly in highly challenging flare scenarios.

图像融合语言控制退化感知注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。