建模人与AI共构知识的动态演化,揭示可持续增长的关键机制。
Dynamics of Human-AI Collective Knowledge on the Web: A Scalable Model and Insights for Sustainable Growth
- 构建可解释的动态模型,模拟知识规模、质量与人机技能的协同演进
- 发现健康增长、反向流动等不同演化模式,量化政策杠杆的影响
- 基于维基百科数据验证模型,揭示大模型兴起后的知识生产失衡
人类与大语言模型(LLMs)如今共同生产并消费网络共享的知识库。这种人机协同知识生态存在正反馈(如加速增长、学习便捷)与系统性风险(如质量下降、能力退化、模型崩溃)。为此,我们提出一个最小化、可解释的动力学模型,刻画知识库规模、质量、模型技能、人类总体技能及查询量的共演化。模型包含两类内容流入(人类、模型),由模型内容准入门控调节;人类两种学习路径(查阅知识库、依赖模型辅助);模型两种训练方式(语料驱动扩展、人类反馈学习)。数值实验揭示了多种增长模式(如健康增长、反向流动、反向学习、振荡),并展示平台策略(门控严格度、模型训练方式、人类学习路径)如何推动系统跨模式跃迁。以文献(PubMed)、代码(GitHub + Copilot)为例,展示不同增长速率与治理规范下的稳态差异。我们将模型拟合至维基百科在预ChatGPT与后ChatGPT时代的知识流数据,发现模型新增占比上升而人类贡献下降,符合模型预测的特定模式。该研究为网络上人机协同知识的可持续发展提供可操作洞见。
原文摘要 · Abstract (English)
Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Such human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and systemic risks (e.g., quality dilution, skill reduction, model collapse). To understand such phenomena, we propose a minimal, interpretable dynamical model of the co-evolution of archive size, archive quality, model (LLM) skill, aggregate human skill, and query volume. The model captures two content inflows (human, LLM) controlled by a gate on LLM-content admissions, two learning pathways for humans (archive study vs. LLM assistance), and two LLM-training modalities (corpus-driven scaling vs. learning from human feedback). Through numerical experiments, we identify different growth regimes (e.g., healthy growth, inverted flow, inverted learning, oscillations), and show how platform and policy levers (gate strictness, LLM training, human learning pathways) shift the system across regime boundaries. Two domain configurations (PubMed, GitHub and Copilot) illustrate contrasting steady states under different growth rates and moderation norms. We also fit the model to Wikipedia's knowledge flow during pre-ChatGPT and post-ChatGPT eras separately. We find a rise in LLM additions with a concurrent decline in human inflow, consistent with a regime identified by the model. Our model and analysis yield actionable insights for sustainable growth of human-AI collective knowledge on the Web.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。