arXiv:2601.21187cs.CVcs.LG2026-01中稿 · ICML

通过子空间级融合,精细注入推理能力而不损失视觉性能。

FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models

  • 基于SVD分解推理模型,按子空间细粒度调控融合强度。
  • 在多个视觉语言推理基准上表现优异,同时保持强视觉能力。
  • 无需标签的自蒸馏训练,适合提升通用视觉语言模型推理力。

通过将大型推理模型(LRM)与视觉语言模型(VLM)高效融合,以增强其推理能力,已成为一个有前景的方向。然而,现有方法通常在粗粒度层级别操作,常导致推理能力增强与视觉能力保留之间的权衡。为此,我们提出FRISM(基于子空间级模型融合的细粒度推理注入),一种基于子空间级模型融合的细粒度推理注入框架。观察到不同SVD子空间对推理和感知的贡献各异,FRISM通过奇异值分解(SVD)分解LRM任务向量,并通过学习自适应调节每个子空间的缩放系数,实现细粒度推理注入。此外,我们引入一种无标签的自蒸馏学习策略,利用通用视觉-语言感知数据集进行双目标优化。大量实验表明,FRISM在持续提升推理能力的同时,大幅保留了模型的视觉能力,在多种视觉语言推理基准上均取得优异表现。

原文摘要 · Abstract (English)

Efficiently enhancing the reasoning capabilities of Vision-Language Models (VLMs) by merging them with Large Reasoning Models (LRMs) has emerged as a promising direction. However, existing methods typically operate at a coarse-grained layer level, which often leads to a trade-off between injecting reasoning capabilities and preserving visual capabilities. To address this limitation, we propose FRISM (Fine-grained Reasoning Injection via Subspace-level model Merging), a fine-grained reasoning injection framework based on subspace-level model merging. Observing that different SVD subspaces contribute differently to reasoning and perception, FRISM decomposes LRM task vectors via Singular Value Decomposition (SVD) and adaptively tunes the scaling coefficients of each subspace through learning to realize fine-grained reasoning injection. Furthermore, we introduce a label-free self-distillation learning strategy with dual-objective optimization using common vision-language perception datasets. Extensive experiments demonstrate that FRISM effectively improves reasoning capabilities while largely preserving the model's visual capabilities by consistently achieving strong performance across diverse visual-language reasoning benchmarks.

视觉语言推理增强模型融合子空间分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。