arXiv:2503.01754cs.CV2025-03被引 4

通过多样化推理路径自蒸馏,提升视觉语言模型的推理能力

SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

  • 用定制提示库生成多样推理问题,分两步推导答案
  • 在五个VQA数据集上显著超越基线模型
  • 适合需要强跨模态推理能力的研究者

推理在各类任务中愈发重要。尽管链式思维提示可有效激发大语言模型的推理能力,但如何发挥视觉语言模型(VLMs)的推理潜力仍具挑战。为此,本文提出一种新型自蒸馏框架以增强模型推理能力。该框架引入多项创新:首先,使用针对视觉推理任务设计的提示库生成多样化的上下文内问题,并采用两步推理流程生成引导性回答;这些回答用于自蒸馏,使模型内化推理过程。此外,还改进模型架构,包括引入干预适配器实现高效参数更新、跨模态跳跃连接促进模态间信息交换,以及集成学习算法融合多条上下文问题产生的多样化推理结果。大量实验表明,该方法在五个VQA数据集上均显著提升基线性能。

原文摘要 · Abstract (English)

Reasoning is increasingly crucial for various tasks. While chain-of-thought prompting enables large language models to leverage reasoning effectively, harnessing the reasoning capabilities of Vision-Language Models (VLMs) remains challenging. To solve this problem, we propose a novel self-distillation framework that enhances the reasoning capabilities of the model. The proposed framework introduces several key innovations. We start by employing a prompt library tailored to visual reasoning tasks to generate diverse in-context questions and utilize a two-step reasoning procedure to derive reasoning-guided responses. These responses are then used for self-distillation, enabling the model to internalize the reasoning process. Additionally, we improve the model architecture with several innovative components, including an intervention adapter for efficient parameter updates, a cross-modal skip connection to facilitate information exchange between modalities, and an ensemble learning algorithm to integrate diverse reasoning from multiple in-context questions. Extensive experiments show that our method significantly improves the baseline performance across five VQA datasets.

视觉语言模型推理增强自蒸馏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。