模型生态百花齐放,评估标准却越来越集中。
Emergent evaluation hubs in a decentralizing large language model ecosystem
- 分析两大生态数据集,发现模型多样性上升但评估基准趋同。
- 前15%节点掌控超80%评估路径,全球评估权威集中度达0.89。
- 适合关注AI评估体系、标准制定与生态治理的研究者。
大型语言模型持续涌现,伴随而来的还有作为共同评价尺度的基准测试。我们探究这两层的集聚模式如何演变:是同步发展还是背道而驰?基于斯坦福基础模型生态系统图谱和Evidently AI基准注册表两个精选代理数据,发现两者呈现互补但相反的趋势。模型创建在国家、组织、模态、许可和访问方式上显著多元化;而基准影响力则呈现集中化趋势:在推断的基准-作者-机构网络中,前15%节点占据超过80%的高介数路径,三个国家贡献了83%的基准输出,推断的全球基准权威基尼系数高达0.89。基于代理的模拟揭示三个机制:新基准的高频引入可降低集中度;快速流入会暂时加剧评估协调难度;对过拟合更强的惩罚影响有限。总体表明,集中的评估影响力充当了协调基础设施,支持标准化、可比性和可复现性,应对模型生产的日益异质性,但也带来路径依赖、选择性可见性和排行榜饱和导致判别力下降等权衡。
原文摘要 · Abstract (English)
Large language models are proliferating, and so are the benchmarks that serve as their common yardsticks. We ask how the agglomeration patterns of these two layers compare: do they evolve in tandem or diverge? Drawing on two curated proxies for the ecosystem, the Stanford Foundation-Model Ecosystem Graph and the Evidently AI benchmark registry, we find complementary but contrasting dynamics. Model creation has broadened across countries and organizations and diversified in modality, licensing, and access. Benchmark influence, by contrast, displays centralizing patterns: in the inferred benchmark-author-institution network, the top 15% of nodes account for over 80% of high-betweenness paths, three countries produce 83% of benchmark outputs, and the global Gini for inferred benchmark authority reaches 0.89. An agent-based simulation highlights three mechanisms: higher entry of new benchmarks reduces concentration; rapid inflows can temporarily complicate coordination in evaluation; and stronger penalties against over-fitting have limited effect. Taken together, these results suggest that concentrated benchmark influence functions as coordination infrastructure that supports standardization, comparability, and reproducibility amid rising heterogeneity in model production, while also introducing trade-offs such as path dependence, selective visibility, and diminishing discriminative power as leaderboards saturate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。