arXiv:2411.08003cs.AIcs.CL2024-11被引 1

大模型生成的恶意文本难以溯源,理论与实证均证明其不可识别。

Can adversarial attacks by large language models be attributed?

  • 将大模型输出视为形式语言,用识别极限理论分析溯源可行性。
  • 即使只有有限样本,多数大模型输出也无法唯一确定来源。
  • 模型数量指数级增长,导致实际溯源成本过高且不现实。

在网络安全攻击和虚假信息传播等对抗性场景中,对大型语言模型(LLM)输出进行溯源面临巨大挑战,且该问题重要性可能持续上升。本文从理论与实证双重视角切入,结合形式语言理论(识别极限)与对不断扩展的LLM生态的数据分析。通过将LLM的可能输出集建模为形式语言,研究有限文本样本能否唯一确定其来源模型。结果表明,在模型能力存在重叠的合理假设下,某些类别的LLM从根本上无法仅凭输出进行识别。我们划分出四种理论可识别性情形:(1) 无限个确定性(离散)LLM语言类不可识别(源于Gold 1967年经典结果);(2) 无限个概率性LLM语言类也不可识别(由确定性情况延伸);(3) 有限个确定性LLM语言类可识别(符合Angluin的可识别准则);(4) 即使是有限个概率性LLM语言类也可能不可识别(本文提出新反例证实此负结果)。此外,我们量化了近年来单个输出对应潜在来源模型数量的爆炸式增长。即使保守假设下——每个开源模型仅微调一个新数据集——候选模型数量每0.5年翻倍;若允许多数据集组合微调,则翻倍时间短至0.28年。这种组合增长,加上跨所有模型和用户进行暴力似然溯源的巨大计算成本,使得全面溯源在实践中完全不可行。

原文摘要 · Abstract (English)

Attributing outputs from Large Language Models (LLMs) in adversarial settings-such as cyberattacks and disinformation campaigns-presents significant challenges that are likely to grow in importance. We approach this attribution problem from both a theoretical and an empirical perspective, drawing on formal language theory (identification in the limit) and data-driven analysis of the expanding LLM ecosystem. By modeling an LLM's set of possible outputs as a formal language, we analyze whether finite samples of text can uniquely pinpoint the originating model. Our results show that, under mild assumptions of overlapping capabilities among models, certain classes of LLMs are fundamentally non-identifiable from their outputs alone. We delineate four regimes of theoretical identifiability: (1) an infinite class of deterministic (discrete) LLM languages is not identifiable (Gold's classical result from 1967); (2) an infinite class of probabilistic LLMs is also not identifiable (by extension of the deterministic case); (3) a finite class of deterministic LLMs is identifiable (consistent with Angluin's tell-tale criterion); and (4) even a finite class of probabilistic LLMs can be non-identifiable (we provide a new counterexample establishing this negative result). Complementing these theoretical insights, we quantify the explosion in the number of plausible model origins (hypothesis space) for a given output in recent years. Even under conservative assumptions-each open-source model fine-tuned on at most one new dataset-the count of distinct candidate models doubles approximately every 0.5 years, and allowing multi-dataset fine-tuning combinations yields doubling times as short as 0.28 years. This combinatorial growth, alongside the extraordinary computational cost of brute-force likelihood attribution across all models and potential users, renders exhaustive attribution infeasible in practice.

大模型安全模型溯源对抗攻击形式语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。