提出视觉定位推理框架,缓解多模态模型的幻觉问题。
Grounded Chain-of-Thought for Multimodal Large Language Models
- 设计逐步定位视觉线索的接地思维链方法
- 在24,022个样本上验证,多数模型一致性不足
- 适合关注视觉推理与可信生成的研究者
尽管取得显著进展,现有多模态大语言模型(MLLMs)仍易产生视觉幻觉,严重阻碍其可信应用。本文从视觉空间推理视角出发,提出一种新学习任务——接地思维链(GCoT),旨在帮助模型逐步识别并定位相关视觉线索,以坐标为依据预测答案。为此,构建了包含5,033张图像和24,022个GCoT示例的多模态接地思维链数据集(MM-GCoT)。同时引入包含答案准确率、定位准确率及答案-定位一致性三项指标的一致性评估体系。在12个先进MLLM上的实验表明:多数模型在一致性评估中表现不佳,存在明显视觉幻觉;且幻觉程度与参数量和通用多模态性能无关,大模型也难幸免。进一步证明,该数据集可有效提升模型的GCoT能力,显著降低不一致回答,并可泛化至开放世界问答与视觉定位等任务。
原文摘要 · Abstract (English)
Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we study this problem from the perspective of visual-spatial reasoning, and propose a new learning task for MLLMs, termed Grounded Chain-of-Thought (GCoT). Different from recent visual CoT studies, which focus more on visual knowledge reasoning, GCoT is keen to helping MLLMs to recognize and ground the relevant visual cues step by step, thereby predicting the correct answer with grounding coordinates as the intuitive basis. To facilitate this task, we also carefully design and construct a dataset called multimodal grounded chain-of-thought (MM-GCoT) consisting of 24,022 GCoT examples for 5,033 images. Besides, a comprehensive consistency evaluation system is also introduced, including the metrics of answer accuracy, grounding accuracy and answer-grounding consistency. We further design and conduct a bunch of experiments on 12 advanced MLLMs, and reveal some notable findings: i. most MLLMs performs poorly on the consistency evaluation, indicating obvious visual hallucination; ii. visual hallucination is not directly related to the parameter size and general multimodal performance, i.e., a larger and stronger MLLM is not less affected by this issue. Lastly, we also demonstrate that the proposed dataset can help existing MLLMs to well cultivate their GCoT capability and reduce the inconsistent answering significantly. Moreover, their GCoT can be also generalized to exiting multimodal tasks, such as open-world QA and REC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。