arXiv:2605.06049cs.CV2026-05中稿 · CVPR

让图像融合听懂不同需求,自动匹配人和机器的偏好。

Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization

论文配图:Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
图 1 · 摘自论文原文
  • 用双扩散模型生成多种融合结果,再通过偏好优化精准控制
  • 在人、视觉语言模型、任务网络间实现精准偏好对齐
  • 适合需要灵活适应不同用户或任务的图像融合场景

红外与可见光图像融合(IVIF)是多模态处理的关键技术,用于整合互补的光谱信息以提升视觉效果并支持下游视觉任务。尽管进展显著,现有方法难以灵活适应多样化的异构需求。实现同时满足人类和机器视觉偏好的自适应融合仍是开放难题。为此,我们提出DPOFusion框架,结合属性对齐的潜在扩散模型(PALDM)与偏好可控的潜在扩散模型(PCLDM),实现面向任务引导、偏好自适应的IVIF。PALDM利用潜在融合先验与联合条件损失,生成具有不同特性的多样化融合候选结果;PCLDM随后通过实例直接偏好优化(IDPO)进行微调,可直接根据异构偏好信号控制最终融合输出。实验表明,该框架不仅在人类、视觉语言模型与任务驱动网络之间实现了精确偏好对齐,还在自适应融合质量与任务导向迁移能力上创下新基准。

原文摘要 · Abstract (English)

As a key technique in multi-modal processing, infrared and visible image fusion (IVIF) plays a crucial role in integrating complementary spectral information for visual enhancement and downstream vision tasks. Despite remarkable progress, existing methods struggle to flexibly accommodate heterogeneous demands. Achieving adaptive fusion that aligns with various preferences from both human and machine vision remains an open and challenging problem. To address this challenge, we propose DPOFusion, a direct preference optimization (DPO) framework integrating the property-aligned latent diffusion model (PALDM) and the preference-controllable latent diffusion model (PCLDM), enabling task-guided, preference-adaptive IVIF for both human and machine vision. The PALDM leverages a latent fusion prior and a joint conditional loss to generate diverse candidate fusion results with various properties. PCLDM is subsequently fine-tuned via instance direct preference optimization (IDPO), enabling direct control of the final fusion results with heterogeneous preference signals. Experimental results demonstrate that our framework not only attains precise preference alignment among humans, vision-language models, and task-driven networks, but also sets a new benchmark for adaptive fusion quality and task-oriented transferability.

图像融合偏好优化扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。