arXiv:2603.03871cs.CV2026-03被引 5

让红外可见光融合图像更符合人眼偏好,提升安防与驾驶辅助效果。

Bridging Human Evaluation to Infrared and Visible Image Fusion

  • 构建首个大规模人类反馈数据集,包含多维度评分与伪影标注。
  • 设计专用奖励模型,使融合结果在主观评价上达到新高度。
  • 适合关注视觉感知优化、人机协同的图像融合研究者。

红外与可见光图像融合(IVIF)通过整合互补模态增强场景感知。当前方法主要聚焦于手工设计损失函数和客观指标,导致融合结果与人类视觉偏好不一致。这一问题因IVIF固有的病态性而加剧,严重限制其在安防监控、驾驶辅助等人类感知场景中的应用。为此,我们提出一种将人类评价融入融合过程的反馈强化框架。针对缺乏以人为中心的评估指标与数据的问题,我们构建了首个大规模人类反馈数据集,包含多维度主观评分与伪影标注,并通过微调的大语言模型进行专家评审增强。基于该数据集,设计领域专用奖励函数并训练奖励模型以量化感知质量。在该奖励引导下,采用分组相对策略优化对融合网络进行微调,实现了优于现有方法的性能,显著提升融合图像与人类审美的一致性。代码已开源。

原文摘要 · Abstract (English)

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is further exacerbated by the ill-posed nature of IVIF, which severely limits its effectiveness in human perceptual environments such as security surveillance and driver assistance systems. To address these limitations, we propose a feedback reinforcement framework that bridges human evaluation to infrared and visible image fusion. To address the lack of human-centric evaluation metrics and data, we introduce the first large-scale human feedback dataset for IVIF, containing multidimensional subjective scores and artifact annotations, and enriched by a fine-tuned large language model with expert review. Based on this dataset, we design a domain-specific reward function and train a reward model to quantify perceptual quality. Guided by this reward, we fine-tune the fusion network through Group Relative Policy Optimization, achieving state-of-the-art performance that better aligns fused images with human aesthetics. Code is available at https://github.com/ALKA-Wind/EVAFusion.

图像融合人类偏好强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。