arXiv:2605.06969cs.CV2026-05被引 1

用大模型模拟人眼判断红外可见光融合图像质量,更精细准确。

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

论文配图:Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
图 1 · 摘自论文原文
  • 让多模态大模型输出连续评分,而非离散等级,提升细微差异辨别力。
  • 基于四个子维度共识构建软标签,反映评价一致性,避免误判。
  • 结合三种监督机制,兼顾图像、方法和场景层级的排序性能。

红外-可见光图像融合(IVIF)旨在将热成像信息与细节空间结构整合到单一图像中以增强感知。现有评估方法过度优化手工设计的无参考统计量和全参考指标,后者将源图像视为伪真实标签。近期基于人类评分的奖励建模采用标量回归,未利用多模态大语言模型(MLLM)的推理能力,也未编码每张图像的感知模糊性,而直接使用离散one-hot监督使质量相近的融合图像被错误划入不同评分等级。为此,我们提出FuScore,利用MLLM模拟人类视觉感知,输出连续质量评分,实现对质量相近图像的细粒度区分。通过四个IVIF特定子维度的一致性构建每张图像的软标签,其锐度反映整体判断的共识程度。进一步引入三重目标:图像级分布监督、源图像对内泰尔斯顿保真度(用于方法级排序),以及跨源图像对泰尔斯顿保真度(用于场景级排序)。大量实验表明,FuScore在与人类视觉偏好相关性上达到当前最佳水平。

原文摘要 · Abstract (English)

Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance perception. However, existing evaluation approaches tend to over-optimize both hand-crafted no-reference statistics and full-reference metrics that treat the source images as pseudo ground truths. Recent IVIF reward-modelling efforts learn from human ratings but use scalar regression on aggregated scores, neither leveraging the reasoning of Multimodal Large Language Models (MLLMs) nor encoding per-image perceptual ambiguity in their supervision, but naively introducing MLLMs with discrete one-hot supervision likewise collapses fused images of similar quality into different rating levels. To address this, we introduce FuScore, which utilizes an MLLM to mimic human visual perception by producing continuous quality score, rather than discrete level predictions, enabling fine-grained discrimination among fused images of similar quality. We exploit the agreement among four IVIF-specific sub-dimensions to construct a per-image soft label whose sharpness reflects how consensual the overall judgment is. We further introduce a tripartite objective combining per-image distributional supervision, within-source-pair Thurstone fidelity for method-level ordering, and cross-source-pair Thurstone fidelity for scene-level ordering across scenes. Extensive experiments demonstrate that FuScore achieves state-of-the-art correlation with human visual preferences.

图像融合大模型评估质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。