arXiv:2603.01416cs.AI2026-03

无需训练即可让视觉语言模型自主搜索,通过融合技术提升性能下限和上限。

Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents

  • 通过跨模态模型融合,将文本搜索代理与基础视觉语言模型结合。
  • 仅用少量校准样本,实现零样本搜索性能基线与更快收敛的高精度峰值。
  • 适合希望快速部署多模态搜索能力且无标注数据的开发者使用。

近期视觉语言模型(VLMs)的发展推动了多模态搜索智能体的出现,这类智能体可主动调用外部搜索工具,并通过多步推理整合检索证据。然而,现有方法通常依赖大规模监督轨迹或昂贵的强化学习,导致训练成本高、不稳定,并存在标准VLM的严重冷启动问题。本文提出一种免训练范式,通过跨模态模型融合赋予VLM自主搜索能力。通过将基于文本的搜索代理与基础VLM融合,我们证明可在不增加多模态训练数据的情况下有效构建多模态搜索能力。为缓解跨模态集成中的参数干扰,引入基于显著性感知的最优大脑融合(OBM)算法,该算法仅需少量校准样本,依据参数对模型损失的影响识别任务关键参数。在多个搜索密集型基准测试(如InfoSeek、MMSearch)上的实验表明:(1) 模型融合作为零样本代理可确保合理的性能基线,其中OBM实现更优的搜索成功率;(2) OBM作为热启动策略显著提升性能上限,收敛速度更快,峰值准确率高于标准VLM初始化。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have motivated the development of multi-modal search agents that can actively invoke external search tools and integrate retrieved evidence through multi-step reasoning. While promising, existing approaches typically rely on large-scale supervised trajectories or expensive reinforcement learning (RL), leading to high training cost, instability, and a severe cold-start problem for standard VLMs. We propose a training-free paradigm to empower VLMs with autonomous search capabilities via cross-modal model merging. By fusing a text-based search agent with a base VLM, we show that multi-modal search capabilities can be effectively composed without any additional multi-modal training data. To mitigate parameter interference during cross-modal integration, we introduce Optimal Brain Merging (OBM), a saliency-aware merging algorithm that identifies task-critical parameters based on their impact on model loss using only a small set of calibration samples. Extensive experiments on search-intensive benchmarks (e.g., InfoSeek, MMSearch) reveal that: (1) Model merging secures a reasonable performance floor as a zero-shot agent, with OBM achieving superior search rates; (2) OBM significantly raises the performance ceiling as a warm-start strategy, achieving faster convergence and higher peak accuracy than standard VLM initialization.

多模态搜索模型融合零样本视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。