arXiv:2508.10367cs.CV2025-08被引 1

用心理物理学方法测大模型的视觉对比敏感度,发现模型差异明显且受提示影响。

Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation

  • 将模型视为整体观察者,通过噪声刺激和二元回答估计对比敏感度函数。
  • 部分模型形状或尺度接近人眼,但无一同时匹配两者,且提示变化显著影响结果。
  • 该方法可预测模型在滤波和对抗样本下的表现,适合评估多模态感知能力。

理解多模态大语言模型(MLLMs)如何处理低层视觉特征,对评估其感知能力至关重要,但尚未系统表征。受人类心理物理学启发,我们提出一种行为方法,将模型视为端到端观察者,通过结构化提示与特定空间频率过滤的噪声刺激进行交互。从二元语义响应中构建心理物理函数,获得对比阈值(及对比敏感度函数,CSF),无需依赖内部激活或分类器代理。结果表明,某些模型在形状或量级上与人类相似,但无一同时匹配二者;且CSF估计对提示表述高度敏感,显示语言鲁棒性有限。最后,我们证明CSF可预测模型在频域滤波与对抗条件下的性能。这些发现揭示了不同MLLM间频率调谐的系统性差异,并确立了CSF估计作为可扩展的多模态感知诊断工具。

原文摘要 · Abstract (English)

Understanding how Multimodal Large Language Models (MLLMs) process low-level visual features is critical for evaluating their perceptual abilities and has not been systematically characterized. Inspired by human psychophysics, we introduce a behavioural method for estimating the Contrast Sensitivity Function (CSF) in MLLMs by treating them as end-to-end observers. Models are queried with structured prompts while viewing noise-based stimuli filtered at specific spatial frequencies. Psychometric functions are derived from the binary verbal responses, and contrast thresholds (and CSFs) are obtained without relying on internal activations or classifier-based proxies. Our results reveal that some models resemble human CSFs in shape or scale, but none capture both. We also find that CSF estimates are highly sensitive to prompt phrasing, indicating limited linguistic robustness. Finally, we show that CSFs predict model performance under frequency-filtered and adversarial conditions. These findings highlight systematic differences in frequency tuning across MLLMs and establish CSF estimation as a scalable diagnostic tool for multimodal perception.

多模态感知评估心理物理对比敏感度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。