arXiv:2504.14871cs.CL2025-04被引 5

即使同数据训练,大模型输出仍可被识别来源,源于训练过程的细微差异。

Natural Fingerprints of Large Language Models

  • 通过控制训练条件,发现模型输出存在可区分的自然指纹
  • 参数量、优化设置、随机种子等微小差异均能留下痕迹
  • 为模型透明性与可靠性研究提供新视角,适合关注AI可解释性的读者

近期研究表明,大型语言模型(LLMs)的输出常能暴露其源模型身份。这虽是模型对训练数据分布建模的自然结果,但这些可识别痕迹也可能反映未经预期的特征,带来公平性与滥用风险。本文进一步表明,即便模型在完全相同的训练数据上训练,其输出仍可区分,说明仅训练动态本身即可留下可识别模式。我们称这些无意的、独特的特征为自然指纹。通过系统控制训练条件,我们发现自然指纹可源于训练过程中的微小差异,如参数规模、优化设置及随机种子。结果表明,训练动态能系统性塑造模型行为,独立于数据与架构,未来在透明性、可靠性和可解释性研究中应予以重视。

原文摘要 · Abstract (English)

Recent studies have shown that the outputs from large language models (LLMs) can often reveal the identity of their source model. While this is a natural consequence of LLMs modeling the distribution of their training data, such identifiable traces may also reflect unintended characteristics with potential implications for fairness and misuse. In this work, we go one step further and show that even when LLMs are trained on exactly the same dataset, their outputs remain distinguishable, suggesting that training dynamics alone can leave recognizable patterns. We refer to these unintended, distinctive characteristics as natural fingerprints. By systematically controlling training conditions, we show that the natural fingerprints can emerge from subtle differences in the training process, such as parameter sizes, optimization settings, and even random seeds. These results suggest that training dynamics can systematically shape model behavior, independent of data or architecture, and should be explicitly considered in future research on transparency, reliability, and interpretability.

大模型指纹训练动态可解释性模型识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。