用数据验证开源对大模型发展的实际影响。
Is Open Source the Future of AI? A Data-Driven Approach
- 收集开源大模型开发数据,分析其改进与贡献趋势。
- 发现开源可提升性能,模型变小且精度损失可控。
- 适合关注AI治理与政策制定的研究者与从业者。
大语言模型在学术界和工业界日益重要,引发对隐私、透明度和滥用的担忧。开源常被视为提升可信度的解决方案,但面临滥用风险、财务激励不足及知识产权问题。相比之下,专有模型凭借企业资源更易实现投资回报。介于完全开源与专有之间的中间路径包括:受许可证保护的开源使用限制、部分开源(开放权重)模型,以及将过时版本开源而保留具有市场价值的版本。当前关于未来模型应处于何种位置的讨论多为观点交锋,缺乏数据支持。本文通过整理大模型开源发展数据,分析其在性能改进、方法创新与修改方面的贡献,旨在为行业专家和政策制定者提供客观依据。研究发现,开源能提升模型性能,呈现模型规模缩小、精度损失可控的趋势,并识别出促进社区参与的积极模式及最受益于开源的架构类型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become central in academia and industry, raising concerns about privacy, transparency, and misuse. A key issue is the trustworthiness of proprietary models, with open-sourcing often proposed as a solution. However, open-sourcing presents challenges, including potential misuse, financial disincentives, and intellectual property concerns. Proprietary models, backed by private sector resources, are better positioned for return on investment. There are also other approaches that lie somewhere on the spectrum between completely open-source and proprietary. These can largely be categorised into open-source usage limitations protected by licensing, partially open-source (open weights) models, hybrid approaches where obsolete model versions are open-sourced, while competitive versions with market value remain proprietary. Currently, discussions on where on the spectrum future models should fall on remains unbacked and mostly opinionated where industry leaders are weighing in on the discussion. In this paper, we present a data-driven approach by compiling data on open-source development of LLMs, and their contributions in terms of improvements, modifications, and methods. Our goal is to avoid supporting either extreme but rather present data that will support future discussions both by industry experts as well as policy makers. Our findings indicate that open-source contributions can enhance model performance, with trends such as reduced model size and manageable accuracy loss. We also identify positive community engagement patterns and architectures that benefit most from open contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。