arXiv:2412.09632cs.CYcs.AI2024-12

通过实验评估英国政府数据对AI训练的实际贡献,揭示哪些数据有用、哪些没被使用。

Methods to Assess the UK Government's Current Role as a Data Provider for AI

  • 用大模型‘遗忘’测试评估政府网站数据重要性
  • 发现政府官网数据对AI有实际影响,但开放数据平台未被使用
  • 方法可复用于其他机构评估自身数据价值

政府通常掌握大量高质量公民与机构数据,英国政府正探索如何更好开放这些数据以促进AI发展。然而,生成式AI训练语料的构成仍是保密信息,难以规划数据共享。为此,本文提出两种方法评估英国政府数据在大型语言模型(LLMs)训练中的作用:一是利用大模型‘未学习’(unlearning)进行消融实验,检验英国政府网站数据对模型性能的影响;二是信息泄露研究,判断大模型是否能识别数据.gov.uk上公开的数据。结果表明,英国政府网站数据在不同主题下对AI具有重要价值,而data.gov.uk数据则未被有效利用。本文为技术报告,详述实验设计、机制与局限,并配套非技术报告在ODI网站发布,总结发现并提出可操作建议。尽管聚焦英国开放数据,但所提方法具备可复现性,可帮助组织评估其数据对AI发展的贡献。

原文摘要 · Abstract (English)

Governments typically collect and steward a vast amount of high-quality data on their citizens and institutions, and the UK government is exploring how it can better publish and provision this data to the benefit of the AI landscape. However, the compositions of generative AI training corpora remain closely guarded secrets, making the planning of data sharing initiatives difficult. To address this, we devise two methods to assess UK government data usage for the training of Large Language Models (LLMs) and 'peek behind the curtain' in order to observe the UK government's current contributions as a data provider for AI. The first method, an ablation study that utilises LLM 'unlearning', seeks to examine the importance of the information held on UK government websites for LLMs and their performance in citizen query tasks. The second method, an information leakage study, seeks to ascertain whether LLMs are aware of the information held in the datasets published on the UK government's open data initiative data$.$gov$.$uk. Our findings indicate that UK government websites are important data sources for AI (heterogenously across subject matters) while data$.$gov$.$uk is not. This paper serves as a technical report, explaining in-depth the designs, mechanics, and limitations of the above experiments. It is accompanied by a complementary non-technical report on the ODI website in which we summarise the experiments and key findings, interpret them, and build a set of actionable recommendations for the UK government to take forward as it seeks to design AI policy. While we focus on UK open government data, we believe that the methods introduced in this paper present a reproducible approach to tackle the opaqueness of AI training corpora and provide organisations a framework to evaluate and maximize their contributions to AI development.

AI训练数据评估政府数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。