arXiv:2601.05600cs.CVcs.CL2026-01ACL被引 11

用场景图引导视觉推理,防止大模型胡说八道。

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

  • 用场景图识别关键实体,主动制造错误视觉信息
  • 构建语言合理但视觉错误的反例对,提升模型准确性
  • 适合需要精准视觉推理的复杂场景应用

多模态大模型在复杂视觉场景中常因缺乏精确的视觉定位而产生推理失真,表现为虚构实体、关系错位、步骤遗漏和过度解释。现有基于偏好的方法依赖文本扰动或答案条件化推理链,容易让模型利用语言先验绕过视觉验证。为此,我们提出SceneAlign框架,利用场景图作为结构化视觉信息,实施可控的结构性干预。通过识别推理关键节点,并采用四种模拟典型定位失败的策略进行扰动,构建出语言上合理但视觉事实错误的硬负例推理链。这些对比样本用于直接偏好优化,引导模型实现细粒度、结构忠实的推理。在七个视觉推理基准上,SceneAlign持续提升答案准确率与推理忠实性,验证了面向定位感知的对齐方法的有效性。

原文摘要 · Abstract (English)

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness frequently manifests as hallucinated entities, mis-grounded relations, skipped steps, and over-specified reasoning. Existing preference-based approaches, typically relying on textual perturbations or answer-conditioned rationales, fail to address this challenge as they allow models to exploit language priors to bypass visual grounding. To address this, we propose SceneAlign, a framework that leverages scene graphs as structured visual information to perform controllable structural interventions. By identifying reasoning-critical nodes and perturbing them through four targeted strategies that mimic typical grounding failures, SceneAlign constructs hard negative rationales that remain linguistically plausible but are grounded in inaccurate visual facts. These contrastive pairs are used in Direct Preference Optimization to steer models toward fine-grained, structure-faithful reasoning. Across seven visual reasoning benchmarks, SceneAlign consistently improves answer accuracy and reasoning faithfulness, highlighting the effectiveness of grounding-aware alignment for multimodal reasoning.

多模态推理场景图视觉定位模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。