发现大模型视觉盲区,提出量化感知边界的新方法
Just Noticeable Difference for Large Multimodal Models
- 提出LMM-JND概念与评测流程,系统评估多任务下模型感知极限
- 构建含48.9万样本的VPA-JND数据集,揭示主流模型性能远低于人类
- 揭示视觉与语言主干设计哲学关联,指导未来模型视觉精度优化
Just noticeable difference(JND)是人类视觉系统可察觉的最小变化,已有数十年研究。尽管近年工作已将该研究拓展至机器视觉,但对大型多模态模型(LMMs)在多种任务和刺激类型下的感知边界仍缺乏系统探索,尤其在当前LMM快速发展的背景下。此外,现有研究未充分考察LMM的感知缺陷,可能带来安全隐患和响应效率低下。本文首次揭示当前LMM存在显著视觉盲区,提出新概念LMM-JND及其确定流程。为揭示与人类视觉一致的任务行为共性,我们深入多个LMM系列,构建大规模数据集VPA-JND,包含21.5万张参考图像及超过48.9万种失真样本,覆盖12类失真类型,以支持LMM-JND研究。实验表明,包括GPT-4o和InternVL2.5系列在内的顶尖LMM在基础对比任务中表现不佳,远低于人类水平。进一步分析发现,视觉与语言主干的设计哲学存在显著相关性,可指导未来模型视觉敏锐度的改进。本研究强调了LMM-JND作为研究LMM的独特视角的重要性,且可预测的LMM-JND对安全性至关重要。代码与数据将在https://github.com/zijianchen98/LMM-JND发布。
原文摘要 · Abstract (English)
Just noticeable difference (JND), the minimum change that the human visual system (HVS) can perceive, has been studied for decades. Although recent work has extended this line of research into machine vision, there has been a scarcity of studies systematically exploring its perceptual boundaries across multiple tasks and stimulus types, particularly in the current era of rapidly advancing large multimodal models (LMMs), where studying the multifaceted capabilities of models has become a mainstream focus. Moreover, the perceptual defects of LMMs are not investigated thoroughly, resulting in potential security issues and suboptimal response efficiency. In this paper, we take an initial attempt and demonstrate that there exist significant visual blind spots in current LMMs. To systemically quantify this characteristic, we propose a new concept, {\bf LMM-JND}, together with its determination pipeline. Targeting uncovering the behavior commonalities in HVS-aligned visual perception tasks, we delve into several LMM families and construct a large-scale dataset, named VPA-JND, which contains 21.5k reference images with over 489k stimuli across 12 distortion types, to facilitate LMM-JND studies. VPA-JND exposes areas where state-of-the-art LMMs, including GPT-4o and the InternVL2.5 series, struggle with basic comparison queries and fall significantly short of human-level visual performance. We further explore the effects of vision and language backbones and find a notable correlation between their design philosophy that may instruct the future refinement of LMMs for their visual acuity. Together, our research underscores the significance of LMM-JND as a unique perspective for studying LMMs, and predictable LMM-JND is crucial for security concerns. This work will be available at https://github.com/zijianchen98/LMM-JND.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。