arXiv:2410.10817cs.CVcs.LG2024-10NeurIPS被引 23

让视觉模型更贴近人类感知,提升多种任务表现。

When Does Perceptual Alignment Benefit Vision Representations?

  • 用人类对图像相似性的判断微调模型,注入感知偏好。
  • 在计数、分割、深度估计等任务上超越原始模型性能。
  • 适合关注人机感知对齐的视觉算法研究者。

人类根据场景布局、主体位置和相机视角等多种视觉属性判断感知相似性。现有视觉模型虽能理解广泛语义抽象,但对这些属性权重分配不当,导致推断与人类感知不一致。尽管感知对齐在图像生成中已带来收益,其在通用视觉任务中的价值仍不明确。本文通过在图像三元组的人类相似性判断上微调前沿模型,并在标准视觉基准上评估其表现。结果表明,对齐人类感知的表示在计数、分割、深度估计、实例检索及检索增强生成等众多下游任务中优于原始骨干网络;同时,在医学影像、3D环境帧等特殊分布外任务中性能也保持稳定。研究说明,向视觉模型引入人类感知知识的归纳偏置,有助于构建更优表示。

原文摘要 · Abstract (English)

Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perception. While vision representations have previously benefited from alignment in contexts like image generation, the utility of perceptually aligned representations in more general-purpose settings remains unclear. Here, we investigate how aligning vision model representations to human perceptual judgments impacts their usability across diverse computer vision tasks. We finetune state-of-the-art models on human similarity judgments for image triplets and evaluate them across standard vision benchmarks. We find that aligning models to perceptual judgments yields representations that improve upon the original backbones across many downstream tasks, including counting, segmentation, depth estimation, instance retrieval, and retrieval-augmented generation. In addition, we find that performance is widely preserved on other tasks, including specialized out-of-distribution domains such as in medical imaging and 3D environment frames. Our results suggest that injecting an inductive bias about human perceptual knowledge into vision models can contribute to better representations.

感知对齐视觉表征微调人类认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。