提出新基准,评估多模态大模型在低级视觉感知中的自我意识水平
Explore the Hallucination on Low-level Perception for MLLMs
- 构建QL-Bench基准,通过问答测试模型对清晰度、光照等低级视觉属性的认知
- 15个模型中多数虽能回答简单问题,但复杂任务下准确率下降且自我反思能力弱
- 适合关注模型可靠性与认知机制的AI研究者,尤其聚焦低级视觉理解方向
多模态大语言模型(MLLMs)在视觉感知与理解方面展现出强大能力,但其幻觉现象严重限制了可靠性,尤其是在低级视觉感知任务中。本文认为幻觉源于模型缺乏显式的自我意识,影响整体性能。为此,我们提出QL-Bench基准,模拟人类对低级视觉的反应,通过视觉问答评估模型在清晰度、光照等低级属性上的自我意识。构建了包含2,990张单图和1,999对图像的LLSAVisionQA数据集,每张图配以开放性问题。对15个MLLM的评估表明,尽管部分模型具备较强低级视觉能力,其自我意识仍较薄弱;相同模型面对简单问题更准确,而复杂问题反而提升自我反思能力。期望该基准推动对MLLM低级视觉自省能力的研究。
原文摘要 · Abstract (English)
The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models also exhibit hallucinations, which limit their reliability as AI systems, especially in tasks involving low-level visual perception and understanding. We believe that hallucinations stem from a lack of explicit self-awareness in these models, which directly impacts their overall performance. In this paper, we aim to define and evaluate the self-awareness of MLLMs in low-level visual perception and understanding tasks. To this end, we present QL-Bench, a benchmark settings to simulate human responses to low-level vision, investigating self-awareness in low-level visual perception through visual question answering related to low-level attributes such as clarity and lighting. Specifically, we construct the LLSAVisionQA dataset, comprising 2,990 single images and 1,999 image pairs, each accompanied by an open-ended question about its low-level features. Through the evaluation of 15 MLLMs, we demonstrate that while some models exhibit robust low-level visual capabilities, their self-awareness remains relatively underdeveloped. Notably, for the same model, simpler questions are often answered more accurately than complex ones. However, self-awareness appears to improve when addressing more challenging questions. We hope that our benchmark will motivate further research, particularly focused on enhancing the self-awareness of MLLMs in tasks involving low-level visual perception and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。