arXiv:2511.17910cs.CL2025-11AAAI被引 4

无需训练,将语言模型的推理能力注入视觉语言模型

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

  • 通过频域干预提取语言模型低频推理特征并注入视觉模型
  • 在多个多步推理任务上超越无训练基线,媲美有监督方法
  • 适合希望低成本提升视觉推理能力的研究者与开发者

链式思维(CoT)推理显著提升了大语言模型(LLMs)的能力,但视觉语言模型(VLMs)因多模态推理数据有限,仍难以胜任多步推理任务。现有方法或需高昂训练成本,或要求架构对齐。本文通过线性人工断层扫描(LAT)实证发现,尽管架构不同,LLMs与VLMs在CoT推理中共享相似的低频潜在表示。基于此,提出L2V-CoT:一种无需训练的潜在干预方法,从LLMs的频域中提取并重采样低频CoT表示,实现维度匹配后在推理阶段注入VLM,增强其推理能力。大量实验表明,该方法持续优于无训练基线,甚至超越部分有监督方法。

原文摘要 · Abstract (English)

Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision-Language Models (VLMs) still struggle with multi-step reasoning tasks due to limited multimodal reasoning data. To bridge this gap, researchers have explored methods to transfer CoT reasoning from LLMs to VLMs. However, existing approaches either need high training costs or require architectural alignment. In this paper, we use Linear Artificial Tomography (LAT) to empirically show that LLMs and VLMs share similar low-frequency latent representations of CoT reasoning despite architectural differences. Based on this insight, we propose L2V-CoT, a novel training-free latent intervention approach that transfers CoT reasoning from LLMs to VLMs. L2V-CoT extracts and resamples low-frequency CoT representations from LLMs in the frequency domain, enabling dimension matching and latent injection into VLMs during inference to enhance reasoning capabilities. Extensive experiments demonstrate that our approach consistently outperforms training-free baselines and even surpasses supervised methods.

多模态推理链式思维零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。