arXiv:2505.10169cs.CVcs.AI2025-05ICCV被引 5

模型因数据集偏差导致跨数据集表现下降超40%,新方法仅用少量参数即可显著提升泛化能力。

Modeling Saliency Dataset Bias

  • 在通用编码器基础上添加少于20个数据集特异性参数,控制多尺度结构等可解释机制
  • 仅用50样本适配新数据,就解决超过75%的跨数据集性能差距
  • 在MIT300、CAT2000、COCO-Freeview上均达新SOTA,适合跨数据集应用

图像显著性预测近年在现有基准上已接近黄金标准表现。然而,由于数据集偏差,跨数据集预测仍具挑战:在某数据集训练的模型应用于另一数据集时,性能下降约40%。令人意外的是,增加数据集多样性无法缓解此差距,近60%的差异源于数据集特有偏差。为此,我们提出一种新型架构,在几乎无偏的编码器-解码器结构基础上,引入少于20个数据集特异性参数,控制多尺度结构、中心偏好和注视扩散等可解释机制。仅调整这些参数即可解决超过75%的泛化差距,且仅需50样本即实现显著提升。该模型在MIT/Tübingen显著性基准(MIT300、CAT2000、COCO-Freeview)上均达到新最优,即使从无关数据集纯泛化也表现优异,若适配对应训练数据则进一步提升。模型还揭示了复杂多尺度效应,涉及绝对与相对尺寸的组合机制。

原文摘要 · Abstract (English)

Recent advances in image-based saliency prediction are approaching gold standard performance levels on existing benchmarks. Despite this success, we show that predicting fixations across multiple saliency datasets remains challenging due to dataset bias. We find a significant performance drop (around 40%) when models trained on one dataset are applied to another. Surprisingly, increasing dataset diversity does not resolve this inter-dataset gap, with close to 60% attributed to dataset-specific biases. To address this remaining generalization gap, we propose a novel architecture extending a mostly dataset-agnostic encoder-decoder structure with fewer than 20 dataset-specific parameters that govern interpretable mechanisms such as multi-scale structure, center bias, and fixation spread. Adapting only these parameters to new data accounts for more than 75% of the generalization gap, with a large fraction of the improvement achieved with as few as 50 samples. Our model sets a new state-of-the-art on all three datasets of the MIT/Tuebingen Saliency Benchmark (MIT300, CAT2000, and COCO-Freeview), even when purely generalizing from unrelated datasets, but with a substantial boost when adapting to the respective training datasets. The model also provides valuable insights into spatial saliency properties, revealing complex multi-scale effects that combine both absolute and relative sizes.

显著性预测数据集偏差泛化能力可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。