首次系统研究3D视觉语言模型幻觉问题,揭示三大成因。
Understanding and Evaluating Hallucinations in 3D Visual Language Models
- 定义3D场景幻觉并分析数据集成因
- 发现物体分布不均、关联强、属性少是主因
- 提出新评估指标,检测模型是否依赖真实视觉
近期提出的3D-LLM(结合点云编码器与大模型)在具身智能和场景理解任务中表现优异,但存在显著幻觉问题,如生成不存在的物体或错误描述物体关系。本文首次对3D-LLM中的幻觉进行系统研究,通过快速评估多个代表性模型,发现其普遍受幻觉影响。我们定义了3D场景中的幻觉,并通过数据集详细分析,揭示三大根源:(1) 物体在数据集中频率分布不均;(2) 物体间存在强相关性;(3) 物体属性多样性有限。此外,提出新的评估指标,包括随机点云对与反向提问评估,用以检验模型生成回答是否基于真实视觉信息且与文本语义一致。
原文摘要 · Abstract (English)
Recently, 3D-LLMs, which combine point-cloud encoders with large models, have been proposed to tackle complex tasks in embodied intelligence and scene understanding. In addition to showing promising results on 3D tasks, we found that they are significantly affected by hallucinations. For instance, they may generate objects that do not exist in the scene or produce incorrect relationships between objects. To investigate this issue, this work presents the first systematic study of hallucinations in 3D-LLMs. We begin by quickly evaluating hallucinations in several representative 3D-LLMs and reveal that they are all significantly affected by hallucinations. We then define hallucinations in 3D scenes and, through a detailed analysis of datasets, uncover the underlying causes of these hallucinations. We find three main causes: (1) Uneven frequency distribution of objects in the dataset. (2) Strong correlations between objects. (3) Limited diversity in object attributes. Additionally, we propose new evaluation metrics for hallucinations, including Random Point Cloud Pair and Opposite Question Evaluations, to assess whether the model generates responses based on visual information and aligns it with the text's meaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。