arXiv:2608.01301cs.CVcs.AI2026-08

提出可模仿人类偏好判断的图像融合评估模型,解决无标准参考时算法排序难题。

Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared-Visible Fusion Assessment

  • 基于人类双选对比实验构建感知偏好模型,支持大规模算法评估
  • 在21场景25方法上实现0.792-0.840的配对准确率,超越传统指标0.163-0.211
  • 结果具反称性与传递性,符合人类判断逻辑,适合科研与工程对比

红外可见光图像融合(IVIF)缺乏理想融合参考,现有算法依赖标量客观指标进行排序,但这些指标常与人类真实偏好不一致。本文提出学习型感知融合度量(LPIFM),将人类双选-等同(A/B/Tie)对比协议转化为可重复、可扩展的代理评估工具。LPIFM联合观察两组源图像和两个融合结果,预测哪个更优或是否等效。训练数据来自新构建的密集偏好语料库,在VIFB基准的21个场景中,对25种融合方法进行全部6,300次无序比较,采用盲化、随机化、双阶段标注并经专家仲裁。在跨场景与跨方法泛化设置下,LPIFM的配对准确率达0.792–0.840,与人工生成的带等同项的Bradley-Terry排名的斯皮尔曼相关系数达0.941–0.977;在完整25方法池中,其准确率优于最强传统指标0.163–0.211。其判断满足候选交换下的反称性,无偏好循环,且完全传递,内部一致性不低于人工小组。代码、模型权重与数据集已公开。

原文摘要 · Abstract (English)

Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize proxies for information transfer, structure, or source similarity. These proxies often disagree with the judgment that ultimately matters: given the same sources, which of two fused results does a human prefer? Direct pairwise comparison is an established protocol for relative subjective assessment, but its cost grows quadratically with the number of algorithms. We present the Learned Perceptual Image Fusion Measure (LPIFM), a source-conditioned model that operationalizes the human A/B/Tie comparison protocol as a repeatable, scalable surrogate. LPIFM jointly observes the two sources and two fused candidates and predicts whether A is better, B is better, or the two are perceptually equivalent. Supervision comes from a new dense preference corpus covering all 6,300 unordered comparisons among 25 fusion methods on the 21 scenes of the VIFB benchmark, labeled under a blinded, randomized, two-stage protocol with expert adjudication. Across scene- and method-generalization settings, LPIFM attains pairwise accuracy of 0.792-0.840 and Spearman correlation of 0.941-0.977 with human-derived tie-aware Bradley-Terry rankings; on full 25-method pools it exceeds the strongest conventional metric by 0.163-0.211 in accuracy. Its verdicts are also antisymmetric under candidate swap, free of preference cycles, and fully transitive, matching or exceeding the internal consistency of the human panel. We publicly release the dataset, model weights, and code. LPIFM offers a practical instrument for human-aligned comparison and ranking of IVIF methods at scale.

图像融合主观评估偏好学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。