arXiv:2501.02964cs.CVcs.AI2025-01被引 10

通过自问自答引导模型关注图像细节,减少幻觉,提升多模态推理能力。

Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild

  • 采用多轮自提问框架,让模型主动聚焦视觉线索。
  • 在细粒度图像任务上使幻觉率降低31.2%。
  • 适合需要高精度视觉理解的轻量级多模态模型研究者。

复杂视觉推理仍是当前挑战。传统方法如思维链(COT)和视觉指令微调虽有效,但二者如何自然结合尚未探索。此外,幻觉与训练成本问题仍待解决。本文提出适用于轻量级多模态大模型的多轮自提问框架——苏格拉底式提问(Socratic Questioning, SQ),通过启发式自提问引导模型关注与目标问题相关的视觉线索,降低幻觉并增强对图像细节的描述能力,从而在复杂视觉推理与问答任务中表现更优。为支持后续研究,我们构建了一个包含1000张细粒度活动图像的多模态小数据集CapQA,用于指令微调与评估。实验表明,所提方法在幻觉评分上提升31.2%,并在多个基准测试中展现出卓越的启发式自提问、零样本视觉推理与幻觉抑制能力。模型与代码将公开。

原文摘要 · Abstract (English)

Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. However, how to organically combine these two methodologies for greater success remains unexplored. Also, issues like hallucinations and high training cost still need to be addressed. In this work, we devise an innovative multi-round training and reasoning framework suitable for lightweight Multimodal Large Language Models (MLLMs). Our self-questioning approach heuristically guides MLLMs to focus on visual clues relevant to the target problem, reducing hallucinations and enhancing the model's ability to describe fine-grained image details. This ultimately enables the model to perform well in complex visual reasoning and question-answering tasks. We have named this framework Socratic Questioning(SQ). To facilitate future research, we create a multimodal mini-dataset named CapQA, which includes 1k images of fine-grained activities, for visual instruction tuning and evaluation, our proposed SQ method leads to a 31.2% improvement in the hallucination score. Our extensive experiments on various benchmarks demonstrate SQ's remarkable capabilities in heuristic self-questioning, zero-shot visual reasoning and hallucination mitigation. Our model and code will be publicly available.

多模态推理自提问幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。