arXiv:2603.22287cs.CVcs.AI2026-03

开源大模型的多模态能力由少数创始事件触发,随后在族谱中快速扩散。

Founder effects shape the evolutionary dynamics of multimodality in open LLM families

  • 通过分析超180万模型的演化谱系,发现多模态能力多源于罕见的创始事件。
  • 2024年后多模态激增,94.5%的视觉语言模型来自已有视觉语言模型的微调。
  • 文本生成模型极少演变为多模态模型,跨模态迁移极弱,适合关注模型演化规律的研究者。

大型语言模型家族发展迅速,但其开放生态中多模态能力的涌现与传播速度尚不明确。基于 Hugging Face 的 ModelBiome AI 生态数据集(超过 180 万条模型元数据与谱系记录),我们量化了多模态能力随时间及父-子关系的演变。跨模态任务在整体生态中早已普遍存在,但在主要开源大模型家族中,多模态直到 2023 年至 2024 年初仍稀少,2024-2025 年急剧上升,以图像-文本视觉语言任务为主。各主流家族首个视觉语言模型(VLM)通常在文本生成模型发布数月后出现,滞后时间从约 1 个月(Gemma)到超过一年,最长达 26 个月(GLM)。谱系条件下的转移率显示跨类型迁移微弱:从文本生成模型微调出的后代成为 VLM 的比例仅为 0.218%。相反,多模态主要在现有 VLM 谱系内扩展:94.5% 的 VLM 子代微调源自 VLM 父模型,仅 4.7% 来自文本生成模型。在模型层面,约 60% 的 VLM 发布为无父模型的新根,其余多为已有 VLM 衍生;创始人集中度分析表明,多模态在谱系内迅速放大并随后多样化。结果表明,多模态通过罕见创始事件进入开源大模型家族,并在后代谱系中快速扩张,形成间断式采纳动态,可能带来多模态能力特有的、受限于迁移的缩放行为。

原文摘要 · Abstract (English)

Large language model (LLM) families are improving rapidly, yet it remains unclear how quickly multimodal capabilities emerge and propagate within open families. Using the ModelBiome AI Ecosystem dataset of Hugging Face model metadata and recorded lineage fields (>1.8x10^6 model entries), we quantify multimodality over time and along recorded parent-to-child relations. Cross-modal tasks are widespread in the broader ecosystem well before they become common within major open LLM families: within these families, multimodality remains rare through 2023 and most of 2024, then increases sharply in 2024-2025 and is dominated by image-text vision-language tasks. Across major families, the first vision-language model (VLM) variants typically appear months after the first text-generation releases, with lags ranging from ~1 month (Gemma) to more than a year for several families and ~26 months for GLM. Lineage-conditioned transition rates show weak cross-type transfer: among fine-tuning edges from text-generation parents, only 0.218% yield VLM descendants. Instead, multimodality expands primarily within existing VLM lineages: 94.5% of VLM-child fine-tuning edges originate from VLM parents, versus 4.7% from text-generation parents. At the model level, most VLM releases appear as new roots without recorded parents (~60%), while the remainder are predominantly VLM-derived; founder concentration analyses indicate rapid within-lineage amplification followed by diversification. Together, these results show that multimodality enters open LLM families through rare founder events and then expands rapidly within their descendant lineages, producing punctuated adoption dynamics that likely induce distinct, transfer-limited scaling behavior for multimodal capabilities.

多模态模型演化谱系分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。