将多模态推理能力注入隐状态,实现更抽象的视觉思维。
OPLD: On-Policy Latent Distillation for Multimodal Reasoning

- 通过策略内蒸馏,让隐状态学习推理过程而非仅对齐视觉特征。
- 在多个基准上超越现有方法,达到当前最优性能。
- 适合研究多模态推理与隐式表示的学者使用。
交错的多模态思维链(CoT)通过引入辅助视觉证据提升视觉推理能力。然而,现有方法受限于外部定义的推理轨迹和视觉操作,难以发展出灵活且抽象的视觉思维。最近,将中间计算内化为连续表示的隐状态推理展现出潜力。但现有视觉-隐状态方法主要通过与压缩的辅助视觉特征对齐来监督隐状态,将其视为视觉观察的代理,而非主动的推理状态。因此,它们虽能捕捉提供证据,却未能充分内化多模态CoT引发的抽象推理过程。本文提出OPLD(On-Policy Latent Distillation),一个简单框架,将特权多模态CoT诱导的推理能力迁移到隐式推理表示中。在多样化的多模态基准上的大量实验表明,OPLD始终优于现有隐状态推理方法,并在多个基准上取得当前最优表现。结果表明,在推理过程层面监督隐状态,比传统的特征层面对齐更有效。
原文摘要 · Abstract (English)
Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。