arXiv:2602.14065cs.AI2026-02中稿 · ICML被引 1

通过推理枢纽解决视觉问答中的知识冲突问题。

REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment

  • 提出推理枢纽概念,聚焦外部证据链接的原子单元。
  • 在多个数据集上提升识别准确率,显著优于现有方法。
  • 适合研究知识融合与视觉问答的学者参考。

知识密集型视觉问答(KI-VQA)常因开放域检索的固有局限而产生严重知识冲突。现有方法缺乏通用的冲突检测机制和模型内约束策略来应对矛盾证据。为此,我们提出基于推理枢纽(Reasoning-Pivot)的REAL框架。不同于侧重内部自推导的推理步骤,推理枢纽是推理链中的原子单元(节点或边),强调知识关联性,并通常依赖外部证据完成推理。依托我们构建的REAL-VQA数据集,该方法通过推理枢纽感知的监督微调(RPA-SFT)训练出可泛化的冲突判别器,同时采用推理枢纽引导解码(RPGD)这一模型内解码策略,利用枢纽实现针对性冲突缓解。在多个数据集上的实验表明,REAL显著提升了判别准确率,验证了以枢纽驱动的冲突解决范式有效性。

原文摘要 · Abstract (English)

Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval. However, existing paradigms face critical limitations due to the lack of generalizable conflict detection and intra-model constraint mechanisms to handle conflicting evidence. To address these challenges, we propose the REAL (Reasoning-Pivot Alignment) framework centered on the novel concept of the Reasoning-Pivot. Distinct from reasoning steps that prioritize internal self-derivation, a reasoning-pivot serves as an atomic unit (node or edge) in the reasoning chain that emphasizes knowledge linkage, and it typically relies on external evidence to complete the reasoning. Supported by our constructed REAL-VQA dataset, our approach integrates Reasoning-Pivot Aware SFT (RPA-SFT) to train a generalizable discriminator by aligning conflicts with pivot extraction, and employs Reasoning-Pivot Guided Decoding (RPGD), an intra-model decoding strategy that leverages these pivots for targeted conflict mitigation. Extensive experiments on diverse datasets demonstrate that REAL significantly enhances discrimination accuracy and achieves superior performance, validating our pivot-driven resolution paradigm.

视觉问答知识冲突推理枢纽多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。