arXiv:2412.09603cs.CV2024-12被引 3

首个大规模测试多模态大模型视觉感知类人行为的基准

Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior

  • 构建包含8.5万样本的HVSBench,覆盖5大视觉领域13类任务
  • 顶尖模型在感知任务上表现中等,远逊于人类表现
  • 适合研究模型与人类视觉系统对齐的AI可解释性方向

尽管多模态大语言模型(MLLMs)在众多视觉任务中表现优异,但其是否具备类人视觉感知行为尚不明确。为此,我们提出HVSBench,首个大规模基准,包含超过85,000个样本,用于评估MLLM与人类视觉系统(HVS)的一致性。该基准涵盖5个关键领域:显著性、数数能力、优先级判断、自由浏览和搜索,共13个类别。全面评估显示存在显著感知差距:即使最先进的MLLM也仅达到中等水平,而人类参与者表现远超所有模型。这凸显了HVSBench的高区分度及发展更类人化AI的迫切需求。我们认为该基准将推动下一代可解释性MLLM的发展。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, the first large-scale benchmark with over 85,000 samples designed to test MLLM alignment with the human visual system (HVS). The benchmark covers 13 categories across 5 key fields: Prominence, Subitizing, Prioritizing, Free-Viewing, and Searching. Our comprehensive evaluation reveals a significant perceptual gap: even state-of-the-art MLLMs achieve only moderate results. In contrast, human participants demonstrate strong performance, significantly outperforming all models. This underscores the high quality of HVSBench and the need for more human-aligned AI. We believe our benchmark will be a critical tool for developing the next generation of explainable MLLMs.

多模态模型视觉感知人类对齐评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。