arXiv:2507.01844cs.CLcs.LG2025-07ACL被引 1

通过分析高概率文本片段,揭示大模型如何复现训练数据。

Low-Perplexity LLM-Generated Sequences and Where To Find Them

  • 用低困惑度序列识别模型生成的高可信文本
  • 发现大量高概率文本无法匹配到训练数据来源
  • 可帮助理解模型对训练数据的复现程度与偏差

随着大语言模型广泛应用,理解其训练数据如何影响输出对透明性、责任性、隐私和公平性至关重要。本文提出一种系统方法,聚焦分析低困惑度序列——即模型生成的高概率文本片段。该流程能可靠地提取跨多个主题的长序列,避免生成退化,并追溯其在训练数据中的来源。令人意外的是,相当一部分低困惑度片段无法映射到训练语料库。对于可匹配的部分,我们量化了其在源文档中的出现分布,揭示了逐字复现的范围与性质,为理解训练数据如何塑造模型行为提供了新路径。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly widespread, understanding how specific training data shapes their outputs is crucial for transparency, accountability, privacy, and fairness. To explore how LLMs leverage and replicate their training data, we introduce a systematic approach centered on analyzing low-perplexity sequences - high-probability text spans generated by the model. Our pipeline reliably extracts such long sequences across diverse topics while avoiding degeneration, then traces them back to their sources in the training data. Surprisingly, we find that a substantial portion of these low-perplexity spans cannot be mapped to the corpus. For those that do match, we quantify the distribution of occurrences across source documents, highlighting the scope and nature of verbatim recall and paving a way toward better understanding of how LLMs training data impacts their behavior.

大模型文本生成训练数据困惑度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。