arXiv:2502.14182cs.CRcs.LG2025-02被引 7

数据投毒研究可推动大模型安全与机制理解

Multi-Faceted Studies on Data Poisoning can Advance LLM Development

  • 从攻击、可信度、机理三方面重新审视数据投毒
  • 能暴露模型隐藏偏见与幻觉,提升鲁棒性
  • 适合关注大模型安全与可解释性的研究者

大语言模型(LLMs)的生命周期远比传统机器学习模型复杂,涉及多个训练阶段、多样数据来源和不同的推理方法。尽管以往关于数据投毒攻击的研究主要聚焦于大模型的安全漏洞,但实际中这类攻击面临诸多挑战:安全的数据采集、严格的数据清洗以及多阶段训练过程,使得污染数据难以注入或可靠地影响模型行为。针对这些问题,本文提出重新思考数据投毒的作用,并主张开展多维度研究,以促进大模型发展。从威胁视角看,实用的数据投毒策略有助于评估和应对大模型的真实安全风险;从可信度视角看,数据投毒可用于揭示并缓解模型中的隐含偏见、有害输出和幻觉问题,从而构建更鲁棒的大模型;从机理视角看,数据投毒能提供关于数据与模型行为之间相互作用的深刻洞察,推动对大模型内在机制的深入理解。

原文摘要 · Abstract (English)

The lifecycle of large language models (LLMs) is far more complex than that of traditional machine learning models, involving multiple training stages, diverse data sources, and varied inference methods. While prior research on data poisoning attacks has primarily focused on the safety vulnerabilities of LLMs, these attacks face significant challenges in practice. Secure data collection, rigorous data cleaning, and the multistage nature of LLM training make it difficult to inject poisoned data or reliably influence LLM behavior as intended. Given these challenges, this position paper proposes rethinking the role of data poisoning and argue that multi-faceted studies on data poisoning can advance LLM development. From a threat perspective, practical strategies for data poisoning attacks can help evaluate and address real safety risks to LLMs. From a trustworthiness perspective, data poisoning can be leveraged to build more robust LLMs by uncovering and mitigating hidden biases, harmful outputs, and hallucinations. Moreover, from a mechanism perspective, data poisoning can provide valuable insights into LLMs, particularly the interplay between data and model behavior, driving a deeper understanding of their underlying mechanisms.

大模型安全数据投毒模型可信度机理研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。