arXiv:2601.14311cs.CRcs.AI2026-01综述被引 4

梳理大模型训练数据的来源、透明度与可追溯性,构建领域分类体系。

Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs

  • 提出数据溯源、透明度与可追溯性的三维框架,涵盖生成、水印等方法
  • 分析95篇论文,提炼出数据标注、偏见测量、隐私保护等核心方法
  • 适合关注模型可解释性与数据合规的研究者与从业者

大规模语言模型已广泛部署,但其训练数据生命周期仍不透明。本文综述过去十年在数据溯源、透明度与可追溯性三个紧密关联维度上的研究,并涵盖偏见与不确定性、数据隐私以及工具技术三大支撑支柱。核心贡献是提出一个分类体系,明确该领域的研究范畴与对应成果。通过对95篇文献的分析,识别出数据生成、水印技术、偏见评估、数据清洗、隐私保护等关键方法,揭示透明度与模糊性之间的固有权衡。

原文摘要 · Abstract (English)

Large language models (LLMs) are deployed at scale, yet their training data life cycle remains opaque. This survey synthesizes research from the past ten years on three tightly coupled axes: (1) data provenance, (2) transparency, and (3) traceability, and three supporting pillars: (4) bias \& uncertainty, (5) data privacy, and (6) tools and techniques that operationalize them. A central contribution is a proposed taxonomy defining the field's domains and listing corresponding artifacts. Through analysis of 95 publications, this work identifies key methodologies concerning data generation, watermarking, bias measurement, data curation, data privacy, and the inherent trade-off between transparency and opacity.

数据溯源模型透明度可追溯性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。