用视觉知识增强推理,让模型能自纠错并解释思考过程。
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
- 通过视觉关系检测提取细粒度视觉知识,提升理解精度。
- 在多个数据集上达到新SOTA,性能媲美ChatGPT-5。
- 引入证据链提示与自我反思机制,实现可解释的自主纠错。
视觉推理旨在解答关于视觉信息的问题。现有方法多依赖预训练视觉语言模型或深度神经网络,但受限于推理过程不可解释,且问题文本存在表述不明确的问题。同时,缺乏细粒度视觉知识也限制了对主体行为的精准理解。为此,我们提出视觉知识驱动的自强化推理框架VIKSER。该框架利用大语言模型的知识蒸馏进行训练,结合视觉关系检测技术提取细粒度视觉知识,并用以重述含模糊表述的问题。我们设计了一种名为证据链(Chain-of-Evidence, CoE)的新提示方法,借助“推理证据”赋予模型可解释的推理能力。同时,集成自我反思技术,使VIKSER能够从错误中学习并持续改进。在多个主流数据集上的实验表明,VIKSER在相关任务中达到新的最先进水平,性能与最新专有模型如ChatGPT-5相当。
原文摘要 · Abstract (English)
Visual reasoning refers to the task of solving questions about visual information. Current visual reasoning methods typically employ pre-trained vision-language model (VLM) strategies or deep neural network approaches. However, existing efforts are constrained by limited reasoning interpretability, while hindering by the phenomenon of underspecification in the question text. Additionally, the absence of fine-grained visual knowledge limits the precise understanding of subject behavior in visual reasoning tasks. To address these issues, we propose VIKSER (Visual Knowledge-Driven Self-Reinforcing Reasoning Framework). Specifically, VIKSER, trained using knowledge distilled from large language models, extracts fine-grained visual knowledge with the assistance of visual relationship detection techniques. Subsequently, VIKSER utilizes fine-grained visual knowledge to paraphrase the question with underspecification. Additionally, we design a novel prompting method called Chain-of-Evidence (CoE), which leverages the power of "evidence for reasoning" to endow VIKSER with interpretable reasoning capabilities. Meanwhile, the integration of self-reflection technology empowers VIKSER with the ability to learn and improve from its mistakes. Experiments conducted on widely used datasets demonstrate that VIKSER achieves new state-of-the-art (SOTA) results in relevant tasks. Moreover, VIKSER achieves performance on par with leading proprietary models, such as the latest ChatGPT-5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。