arXiv:2510.10161cs.CLcs.AI2025-10综述被引 6

梳理大模型信息溯源方法,提升生成内容的可信度

Large Language Model Sourcing: A Survey

  • 从模型、结构、训练数据到外部数据四维度系统梳理溯源方法
  • 提出前因与后验双范式分类体系,覆盖主动嵌入与被动推断两类技术
  • 适合关注AI可信性、合规性及可解释性的研究者与开发者

由于大语言模型(LLMs)具有黑箱特性且生成内容高度逼真,幻觉、偏见、不公平及版权侵权等问题日益突出。在此背景下,多视角信息溯源至关重要。本文围绕四个相互关联的维度——模型溯源、模型结构溯源、训练数据溯源和外部数据溯源,开展系统性调查。进一步提出一种统一的双范式分类体系,将现有溯源方法分为基于先验(主动可追溯性嵌入)和基于后验(事后推理)两类。跨维度的可追溯性显著提升了大模型在真实应用中的透明度、问责性与可信度。

原文摘要 · Abstract (English)

Due to the black-box nature of large language models (LLMs) and the realism of their generated content, issues such as hallucinations, bias, unfairness, and copyright infringement have become significant. In this context, sourcing information from multiple perspectives is essential. This survey presents a systematic investigation organized around four interrelated dimensions: Model Sourcing, Model Structure Sourcing, Training Data Sourcing, and External Data Sourcing. Moreover, a unified dual-paradigm taxonomy is proposed that classifies existing sourcing methods into prior-based (proactive traceability embedding) and posterior-based (retrospective inference) approaches. Traceability across these dimensions enhances the transparency, accountability, and trustworthiness of LLMs deployment in real-world applications.

大模型信息溯源可信AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。