通过约束提示减少视觉语言模型的空间关系幻觉
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
- 引入双向与传递性约束,确保物体间关系一致
- 在三个数据集上提升空间关系预测准确率
- 适合研究多模态推理与幻觉抑制的学者
空间关系幻觉是大视觉语言模型(LVLMs)面临的持续挑战,常导致对图像中物体位置和空间配置的错误预测。为解决此问题,我们提出一种约束感知提示框架,引入两类约束:(1) 双向约束,确保成对物体关系的一致性;(2) 传递性约束,强制多物体间关系的依赖性。通过融入这些约束,LVLMs 能生成更空间一致且连贯的输出。我们在三个常用空间关系数据集上评估该方法,结果优于现有方法。此外,对多种双向关系分析选择与传递性参考项的系统性分析,揭示了该方法在融合约束以缓解空间关系幻觉方面的更大潜力。
原文摘要 · Abstract (English)
Spatial relation hallucinations pose a persistent challenge in large vision-language models (LVLMs), leading to generate incorrect predictions about object positions and spatial configurations within an image. To address this issue, we propose a constraint-aware prompting framework designed to reduce spatial relation hallucinations. Specifically, we introduce two types of constraints: (1) bidirectional constraint, which ensures consistency in pairwise object relations, and (2) transitivity constraint, which enforces relational dependence across multiple objects. By incorporating these constraints, LVLMs can produce more spatially coherent and consistent outputs. We evaluate our method on three widely-used spatial relation datasets, demonstrating performance improvements over existing approaches. Additionally, a systematic analysis of various bidirectional relation analysis choices and transitivity reference selections highlights greater possibilities of our methods in incorporating constraints to mitigate spatial relation hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。