评测大模型对物体未完成感知的理解能力,发现日语提示下性能反而下降。
Bridging Perception and Language: A Systematic Benchmark for LVLMs' Understanding of Amodal Completion Reports
- 基于本体论构建系统性评测基准,分类物体缺失感知任务
- 多数大模型整体表现接近人类,但特定物体类别准确率差异显著
- 日语提示下部分模型在有图时表现劣于无图,暴露语言能力短板
开发大型视觉语言模型(LVLMs)的核心目标之一是让其能辅助人类完成多模态任务,包括理解感知体验的描述。其中关键现象是无模态补全——即使物体部分被遮挡,人们仍能感知其完整存在。尽管已有研究评估计算机视觉算法对遮挡区域的检测与重建能力,但对LVLMs在无模态补全相关文本上的推理能力仍缺乏探索。为此,我们基于基础形式本体论构建了一个系统性评测基准,实现对无模态补全的结构化分类。结果显示,虽然多数LVLMs在整体上达到人类水平,但在某些物体类别上表现差异明显。值得注意的是,在特定类别中,部分LLaVA-NeXT变体和Claude 3.5 Sonnet在原始图像上的准确率反而低于无视觉内容的空白刺激。这种反常现象仅出现在日语提示下,暗示这些模型在日语语言理解方面存在缺陷。
原文摘要 · Abstract (English)
One of the main objectives in developing large vision-language models (LVLMs) is to engineer systems that can assist humans with multimodal tasks, including interpreting descriptions of perceptual experiences. A central phenomenon in this context is amodal completion, in which people perceive objects even when parts of those objects are hidden. Although numerous studies have assessed whether computer-vision algorithms can detect or reconstruct occluded regions, the inferential abilities of LVLMs on texts related to amodal completion remain unexplored. To address this gap, we constructed a benchmark grounded in Basic Formal Ontology to achieve a systematic classification of amodal completion. Our results indicate that while many LVLMs achieve human-comparable performance overall, their accuracy diverges for certain types of objects being completed. Notably, in certain categories, some LLaVA-NeXT variants and Claude 3.5 Sonnet exhibit lower accuracy on original images compared to blank stimuli lacking visual content. Intriguingly, this disparity emerges only under Japanese prompting, suggesting a deficiency in Japanese-specific linguistic competence among these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。