无需训练,统一解决多视觉语言模型的选择、适配与集成问题。
One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling

- 基于自适应最优传输,从多个冻结模型中学习目标域的共识样本-类别结构。
- 在跨域场景下,显著提升模型选择准确性、适配稳定性和集成鲁棒性。
- 适用于无标签目标数据的多模型部署,尤其适合医学、遥感等专业领域。
视觉语言模型(VLM)可实现基于语义描述的视觉识别,对缺乏标注的目标尤为有用。现有方法通常先选单一VLM再进行适配,隐含假设所选模型已适配目标域。但在真实跨域场景中,可能存在多个通用或领域专用的候选VLM,却无实例级标签用于甄别可靠模型。因此需同时解决模型选择、目标适配与预测集成。本文提出One Stone, Three Birds(OSTB),从系统级多VLM视角重构该问题。核心观察是三者均依赖同一潜在结构:目标集中的可信样本-类别关系。不同VLM虽有迁移偏差并产生冲突预测,但其输出可互补地推断该结构。我们提出一种无需训练的自适应最优传输框架,给定一组冻结候选VLM,OSTB估计一个共识样本-类别运输计划,且不更新模型参数。该学习到的运输结构被复用于所有部署目标:模型选择通过共识计划诱导的语义与视觉可靠性排序;目标适配通过条件于运输的视觉分类器拟合实现;集成则通过可靠性感知的概率融合完成。在自然图像、遥感与医学病理基准上的大量实验表明,当候选池异质时,OSTB显著提升模型排序、适配稳定性与集成鲁棒性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) enable visual recognition from semantic class descriptions, which makes them attractive when target annotations are scarce or unavailable. Most deployment pipelines, however, first choose a single VLM and then adapt that model to the unlabeled target set. This single-backbone paradigm hides a critical assumption: the selected VLM is already compatible with the target domain. In realistic cross-domain deployment, several general-purpose and domain-specialized VLMs may be plausible, yet no instance-level target labels are available to identify the reliable ones. Deployment therefore requires a coupled solution for model selection, target adaptation, and prediction integration. We revisit this problem from a system-level multi-VLM perspective. Our central observation is that the three decisions above depend on the same latent object: a trustworthy sample-class structure in the target set. Different VLMs may encode different transfer biases and produce conflicting predictions, but their outputs can still provide complementary evidence for estimating this structure. We propose One Stone, Three Birds, a training-free framework based on self-adaptive optimal transport. Given a pool of frozen candidate VLMs, OSTB estimates a consensus sample-to-class transport plan without updating VLM parameters. The learned transport structure is then reused for all deployment objectives: model selection is performed by ranking the combined semantic and visual reliability induced by the consensus plan; target adaptation is obtained by fitting transport-conditioned visual classifiers; and ensembling is implemented through reliability-aware probabilistic integration. Extensive experiments on natural-image, remote-sensing, and medical-pathology benchmarks show that OSTB improves model ranking, adaptation stability, and ensemble robustness under heterogeneous candidate pools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。