arXiv:2608.25663cs.IRcs.AI2026-08

大模型需引用数据来源,但现有方法难以实现可信追溯与公正署名。

Data Citation for Large Language Models: A Challenge

  • 提出三大研究方向:训练数据溯源、推理时数据引用、知识图谱三元组引用
  • 强调需在数据粒度和固定性上精准引用,确保结果可验证
  • 适合关注大模型可解释性与数据伦理的研究者

大型语言模型日益成为信息获取的主要中介,学界开始关注其输出是否应标注来源。现有研究将引用视为验证工具,仅适用于文本文档。学术引用还承担着赋权与溯源功能,对数据同样适用。本文指出,大模型的数据引用是一个开放挑战,不同于文档级引用定位,且更难解决。核心问题在于:如何让模型引用数据,以保证输出可验证、溯源可追踪,并使数据创造者与整理者获得合理认可。为此,论文提出三个研究方向:训练数据归属需将影响估计转化为对语料库的引用;推理时的数据引用需在合适粒度和固定性下识别数据集、子集及查询结果;知识图谱事实引用需明确定义对单个三元组的引用含义及信用传播路径。上述进展依赖数据库、信息检索、知识表示与人工智能领域的协同合作。

原文摘要 · Abstract (English)

Large language models increasingly mediate access to information, and a growing body of work asks whether they cite the sources behind their outputs. That work treats citation as a verification device and applies it to textual documents. Scholarly citation serves two further functions, credit and provenance, and it applies to data as much as to text. This paper argues that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve. We ask how such models should cite data so that outputs stay verifiable, provenance stays traceable, and credit reaches data creators and curators. We set out three research directions. Training data attribution has to turn influence estimates into references for corpora absorbed into model parameters. Data citation at inference time has to identify datasets, subsets, and query results at the right granularity and fixity. Citing knowledge graph facts has to define what a reference to a single triple denotes and how credit propagates along provenance. Progress on all three depends on joint work across the database, information retrieval, knowledge representation, and artificial intelligence communities.

大模型数据引用可追溯性知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。