通过视觉编码器内部动态,精准预测模型在特定任务上的表现。
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
- 用层级导通性表征任务,捕捉模型内部功能块激活模式。
- 提出非对称度量DCD,实现14.7%的NDCG@5提升,优于现有方法。
- 适合少样本场景下快速筛选最优视觉语言模型,节省计算资源。
尽管开源视觉语言模型(VLMs)数量激增,但在特定下游任务中选择最优预训练模型仍具挑战。由于计算资源和少样本数据限制,全面评估往往不可行。现有方法存在缺陷:或依赖大量数据代理,或使用对称文本描述,忽略迁移能力的方向性和模型特异性。为此,我们提出一种基于视觉编码器内部功能动态的模型选择框架。通过逐层导通性表征任务,并利用熵正则化对齐生成目标条件下的块重要性分布。在此基础上,引入方向性导通性差异(DCD),一种非对称度量,量化源任务覆盖目标任务关键功能块的有效性。该方法可在无需直接推理的前提下,通过聚合源任务排名预测目标模型排序。在21个数据集上对48个VLMs的实验表明,本方法显著优于当前最佳基线,NDCG@5提升达14.7%。
原文摘要 · Abstract (English)
While open sourced Vision-Language Models (VLMs) have proliferated, selecting the optimal pretrained model for a specific downstream task remains challenging. Exhaustive evaluation is often infeasible due to computational constraints and data limitations in few shot scenarios. Existing selection methods fail to fully address this: they either rely on data-intensive proxies or use symmetric textual descriptors that neglect the inherently directional and model-specific nature of transferability. To address this problem, we propose a framework that grounds model selection in the internal functional dynamics of the visual encoder. Our approach represents each task via layer wise conductance and derives a target-conditioned block importance distribution through entropy regularized alignment. Building on this, we introduce Directional Conductance Divergence (DCD), an asymmetric metric that quantifies how effectively a source task covers the target's salient functional blocks. This allows for predicting target model rankings by aggregating source task ranks without direct inference. Experimental results on 48 VLMs across 21 datasets demonstrate that our method outperforms state-of-the-art baselines, achieving a 14.7% improvement in NDCG@5 over SWAB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。