arXiv:2605.03885cs.CVcs.LG2026-05

提出新方法提升眼动密度估计精度,让视觉注意力基准更可靠。

Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking

论文配图:Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking
图 1 · 摘自论文原文
  • 用自适应带宽核密度与混合模型结合,优化每张图的密度估计。
  • 在多个基准上提升5-15%对数似然,关键图像改进超25%。
  • 适合做模型失败分析、逆向评估的科研人员参考使用。

基于人类眼动数据估算的实证注视密度是显著性基准的核心。它们直接影响基准结论、排行榜排名、失败案例分析及对人类视觉行为的科学判断。然而,标准的固定带宽各向同性高斯核密度估计(KDE)几十年来未变。随着研究转向样本级评估(如失败案例分析、逆向基准、逐图像模型比较),精确的逐图像密度估计变得至关重要。本文提出一种原理严谨的混合模型,结合自适应带宽KDE(基于Abramson方法)、中心偏差项和均匀成分,并融合前沿显著性模型,以捕捉不同空间与语义层面的观察者一致性;通过留一被试交叉验证对每张图像优化所有参数。该方法在多个基准上显著提升观察者间一致性,图像级对数似然中位数提高5%-15%,AUC最高提升2个百分点;对最相关的失败案例图像,改进超过25%。利用这些优化后的估计,我们识别并分析了当前最优显著性模型的残余失败案例,证明模型仍有巨大提升空间。更广泛地,研究强调实证注视密度不应视为固定真值,而应作为随方法进步持续演进的估计量。

原文摘要 · Abstract (English)

Empirical fixation densities, spatial distributions estimated from human eye-tracking data, are foundational to saliency benchmarking. They directly shape benchmark conclusions, leaderboard rankings, failure case analyses, and scientific claims about human visual behavior. Yet the standard estimation method, fixed-bandwidth isotropic Gaussian KDE, has gone essentially unchanged for decades. This matters now more than ever: as the field shifts toward sample-level evaluation (failure case analysis, inverse benchmarking, per-image model comparison), reliable per-image density estimates become critical. We propose a principled mixture model that combines an adaptive-bandwidth KDE based on Abramson's method, center bias and uniform components, and a state-of-the-art saliency model, to capture different spatial and semantic types of interobserver consistency, and optimize all parameters per image via leave-one-subject-out cross-validation. Our method yields substantially higher interobserver consistency estimates across multiple benchmarks, with median per-image gains of 5-15% in log-likelihood and up to 2 percentage points in AUC. For the most affected images -- precisely those most relevant to failure case analysis -- improvements exceed 25%. We leverage these improved estimates to identify and analyze remaining failure cases of state-of-the-art saliency models, demonstrating that significant headroom for model improvement remains. More broadly, our findings highlight that empirical fixation densities should not be treated as fixed ground truths but as evolving estimates that improve with better methodology.

显著性模型眼动数据密度估计基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。