arXiv:2608.00013cs.CLcs.AI2026-08

用文本能力分数预测视觉语言模型性能,让选基座模型更科学。

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

  • 基于文本能力得分构建跨模型族的可迁移性预测框架。
  • 准确预测从80亿到720亿参数模型的性能表现与训练轨迹。
  • 揭示文本基准陷阱、基座模型优势及不同模型家族差异。

选择合适的大型语言模型(LLM)作为视觉语言模型(VLM)的基座,是构建VLM最关键的决策,但目前仍缺乏系统方法:基于算力的扩展定律无法跨模型族推广,且缺乏训练前直接预测VLM性能的框架。本文提出首个跨家族的基于能力的多模态扩展定律(Capability-Driven Multimodal Scaling Law),通过主成分分析(PCA)提取的低维能力分数 $S$,建模VLM性能与 $S$ 的关系,包含每类基座的迁移率和数据扩展效率吸收率。我们基于34个来自7个模型家族的LLM,在严格控制条件下训练了超过150个VLM。在200多个文本与50多个多模态基准上验证表明,该定律能准确外推80亿至720亿参数模型的迁移率,高保真预测完整训练轨迹,并泛化至完全未见的模型族。分析还发现:某些文本基准与多模态性能负相关,暴露基准“作弊”行为;基座模型比指令微调版本更适合作为VLM基座,因其吸收率更高、数据扩展衰减更低;不同模型家族在迁移-吸收空间中占据不同位置。该框架将基座选择从耗时的试错变为可量化的理性决策。代码与数据见https://github.com/wangq-dev/CDMScaling。

原文摘要 · Abstract (English)

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at https://github.com/wangq-dev/CDMScaling.

视觉语言模型扩展定律模型选择能力评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。