arXiv:2510.01645cs.CRcs.AI2025-10被引 18

LLM隐私风险远超数据记忆,需关注全链条新威胁

Position: Privacy Is Not Just Memorization!

  • 构建从数据收集到部署的全生命周期隐私风险分类体系
  • 分析1322篇论文发现:记忆问题受关注度远高于实际危害
  • 呼吁跨学科研究,应对智能体与推理攻击等新兴威胁

大型语言模型(LLM)的隐私讨论长期聚焦于训练数据的逐字记忆,而诸多更紧迫且可扩展的隐私威胁却未被充分关注。本文指出,LLM系统的隐私风险远超数据提取,涵盖数据采集实践、推理时上下文泄露、自主代理能力,以及通过深度推理攻击实现的监控技术民主化。我们提出了一个覆盖LLM全生命周期的隐私风险综合分类体系,并通过案例研究证明现有隐私框架无法应对这些多维度威胁。对2016—2025年顶级会议发表的1,322篇AI/ML隐私论文进行纵向分析显示,尽管记忆问题在技术研究中备受关注,但最严重的隐私危害恰恰出现在当前技术手段难以应对的领域,且未来路径尚不明确。本文呼吁研究界根本性转变视角,超越现有技术方案的狭隘焦点,采用跨学科方法应对这些新兴威胁的社会技术本质。

原文摘要 · Abstract (English)

The discourse on privacy risks in Large Language Models (LLMs) has disproportionately focused on verbatim memorization of training data, while a constellation of more immediate and scalable privacy threats remain underexplored. This position paper argues that the privacy landscape of LLM systems extends far beyond training data extraction, encompassing risks from data collection practices, inference-time context leakage, autonomous agent capabilities, and the democratization of surveillance through deep inference attacks. We present a comprehensive taxonomy of privacy risks across the LLM lifecycle -- from data collection through deployment -- and demonstrate through case studies how current privacy frameworks fail to address these multifaceted threats. Through a longitudinal analysis of 1,322 AI/ML privacy papers published at leading conferences over the past decade (2016--2025), we reveal that while memorization receives outsized attention in technical research, the most pressing privacy harms lie elsewhere, where current technical approaches offer little traction and viable paths forward remain unclear. We call for a fundamental shift in how the research community approaches LLM privacy, moving beyond the narrow focus of current technical solutions and embracing interdisciplinary approaches that address the sociotechnical nature of these emerging threats.

隐私风险LLM安全社会技术综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。