用大模型模拟人看图习惯,融合红外与可见光图像更符合人类视觉。
Infrared and Visible Image Fusion with Hierarchical Human Perception
- 用大视觉语言模型生成人类关注问题的答案,作为融合依据。
- 融合结果在信息保留和人眼感知上均优于现有方法。
- 适合需要真实视觉体验的安防、医疗成像场景。
图像融合将多源图像整合为一幅包含互补信息的图像。现有方法以像素强度、纹理及高层视觉任务信息为标准来决定信息保留,缺乏对人类感知的增强。本文提出一种名为分层感知融合(HPFusion)的图像融合方法,利用大视觉语言模型引入分层人类语义先验,保留符合人类视觉系统的互补信息。我们设计了人类观察图像对时关注的多个问题,通过大视觉语言模型根据图像生成答案。将答案文本编码输入融合网络,优化目标是使融合图像的人类语义分布更接近源图像,探索人类感知域内的互补信息。大量实验表明,本方法在信息保留和人类视觉增强方面均能取得高质量融合结果。
原文摘要 · Abstract (English)
Image fusion combines images from multiple domains into one image, containing complementary information from source domains. Existing methods take pixel intensity, texture and high-level vision task information as the standards to determine preservation of information, lacking enhancement for human perception. We introduce an image fusion method, Hierarchical Perception Fusion (HPFusion), which leverages Large Vision-Language Model to incorporate hierarchical human semantic priors, preserving complementary information that satisfies human visual system. We propose multiple questions that humans focus on when viewing an image pair, and answers are generated via the Large Vision-Language Model according to images. The texts of answers are encoded into the fusion network, and the optimization also aims to guide the human semantic distribution of the fused image more similarly to source images, exploring complementary information within the human perception domain. Extensive experiments demonstrate our HPFusoin can achieve high-quality fusion results both for information preservation and human visual enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。