剖析主流大模型透明度,揭示开源标签背后的真相
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
- 从开源与开权重双视角评估100+大模型透明度
- 多数标称开源模型仍缺训练数据、代码及碳排放等关键信息
- 警示
尽管人工智能开源讨论日益增多,但对顶尖大语言模型(LLMs)的透明度与可访问性研究仍显不足。本文基于开源倡议(OSI)最新定义,系统分析过去五年内包括ChatGPT、DeepSeek、LLaMA在内的百余款主流大模型,从开源与开权重两个维度评估其透明度。研究发现,即便部分模型被标为开源,也往往未公开训练数据、代码、权重可访问性及碳排放等关键信息。此类“伪开源”现象导致模型难以复现、偏见难消解、领域适配受限。本研究首次构建双重视角框架,揭示当前大模型开放实践中的深层次问题,呼吁推动更负责任、可持续的AI发展。
原文摘要 · Abstract (English)
Despite increasing discussions on open-source Artificial Intelligence (AI), existing research lacks a discussion on the transparency and accessibility of state-of-the-art (SoTA) Large Language Models (LLMs). The Open Source Initiative (OSI) has recently released its first formal definition of open-source software. This definition, when combined with standard dictionary definitions and the sparse published literature, provide an initial framework to support broader accessibility to AI models such as LLMs, but more work is essential to capture the unique dynamics of openness in AI. In addition, concerns about open-washing, where models claim openness but lack full transparency, has been raised, which limits the reproducibility, bias mitigation, and domain adaptation of these models. In this context, our study critically analyzes SoTA LLMs from the last five years, including ChatGPT, DeepSeek, LLaMA, and others, to assess their adherence to transparency standards and the implications of partial openness. Specifically, we examine transparency and accessibility from two perspectives: open-source vs. open-weight models. Our findings reveal that while some models are labeled as open-source, this does not necessarily mean they are fully open-sourced. Even in the best cases, open-source models often do not report model training data, and code as well as key metrics, such as weight accessibility, and carbon emissions. To the best of our knowledge, this is the first study that systematically examines the transparency and accessibility of over 100 different SoTA LLMs through the dual lens of open-source and open-weight models. The findings open avenues for further research and call for responsible and sustainable AI practices to ensure greater transparency, accountability, and ethical deployment of these models.(DeepSeek transparency, ChatGPT accessibility, open source, DeepSeek open source)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。