分析186万模型,发现大模型生态像生物进化,兄弟模型比父子更像。
Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- 用基因相似性分析模型家族,发现兄弟模型比父辈更相似
- 许可证从商业限制转向宽松开放,常违反上游条款
- 模型卡变短变模板化,语言能力逐渐只支持英文
本文分析了Hugging Face平台上186万模型的细调谱系。通过构建模型家族树,研究发现细调线谱庞大且结构各异。借鉴进化生物学视角,利用元数据与模型卡片衡量模型间的遗传相似性与性状变异。结果表明,同一家族模型具有显著的家族相似性,但其演化模式不同于无性繁殖:突变速度快且定向,导致兄弟模型间相似度高于父代与子代。进一步分析显示,许可证倾向从限制性商业许可向宽松或著佐许可演变,常违背上游许可条款;模型从多语言支持转向仅支持英文;模型卡长度缩短,日益采用模板化和自动生成内容。该研究为理解开放机器学习生态提供了实证基础,表明生态学方法可揭示模型演化新规律。
原文摘要 · Abstract (English)
Many have observed that the development and deployment of generative machine learning (ML) and artificial intelligence (AI) models follow a distinctive pattern in which pre-trained models are adapted and fine-tuned for specific downstream tasks. However, there is limited empirical work that examines the structure of these interactions. This paper analyzes 1.86 million models on Hugging Face, a leading peer production platform for model development. Our study of model family trees -- networks that connect fine-tuned models to their base or parent -- reveals sprawling fine-tuning lineages that vary widely in size and structure. Using an evolutionary biology lens to study ML models, we use model metadata and model cards to measure the genetic similarity and mutation of traits over model families. We find that models tend to exhibit a family resemblance, meaning their genetic markers and traits exhibit more overlap when they belong to the same model family. However, these similarities depart in certain ways from standard models of asexual reproduction, because mutations are fast and directed, such that two `sibling' models tend to exhibit more similarity than parent/child pairs. Further analysis of the directional drifts of these mutations reveals qualitative insights about the open machine learning ecosystem: Licenses counter-intuitively drift from restrictive, commercial licenses towards permissive or copyleft licenses, often in violation of upstream license's terms; models evolve from multi-lingual compatibility towards english-only compatibility; and model cards reduce in length and standardize by turning, more often, to templates and automatically generated text. Overall, this work takes a step toward an empirically grounded understanding of model fine-tuning and suggests that ecological models and methods can yield novel scientific insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。