让视觉语言模型看懂错觉图,只需缩小图像分辨率。
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
- 将图像缩至32-128像素,激活模型对错觉的感知能力
- 在112张错觉图上,准确率从0-5.36%提升至99%以上
- 适合关注模型鲁棒性与人机认知差距的研究者
视觉语言模型(VLMs)在语义任务中表现优异,但在识别光学错觉或生成图像中的隐藏内容方面表现极差,即使通过明确提示也难以改善。我们构建了HC-Bench基准,包含112张含隐藏文字、物体和错觉的图像,发现主流VLMs准确率仅为0-5.36%。人类能本能地通过视觉调整(如缩放)识别这些内容,而模型因过度依赖高层语义而失败。我们提出SemVink(语义视觉思维),仅通过将图像缩小至32-128像素,即可实现超过99%的准确率,有效去除冗余视觉噪声。这揭示了当前VLM架构的核心缺陷:忽视低层视觉操作,导致真实场景下缺乏鲁棒性。本研究呼吁发展融合多尺度处理的混合模型,推动计算视觉与人类认知的融合,适用于医疗影像、安全检测等关键领域。
原文摘要 · Abstract (English)
Vision-language models (VLMs) excel in semantic tasks but falter at a core human capability: detecting hidden content in optical illusions or AI-generated images through perceptual adjustments like zooming. We introduce HC-Bench, a benchmark of 112 images with hidden text, objects, and illusions, revealing that leading VLMs achieve near-zero accuracy (0-5.36%)-even with explicit prompting. Humans resolve such ambiguities instinctively, yet VLMs fail due to an overreliance on high-level semantics. Strikingly, we propose SemVink (Semantic Visual Thinking) by simply scaling images to low resolutions (32-128 pixels), which unlocks >99% accuracy by eliminating redundant visual noise. This exposes a critical architectural flaw: VLMs prioritize abstract reasoning over low-level visual operations crucial for real-world robustness. Our work urges a shift toward hybrid models integrating multi-scale processing, bridging the gap between computational vision and human cognition for applications in medical imaging, security, and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。