arXiv:2506.15674cs.CLcs.AI2025-06EMNLP被引 35

大模型推理过程会泄露用户隐私,越仔细思考越容易暴露信息。

Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers

  • 通过提示注入可从模型推理链中提取敏感数据
  • 增加推理步骤会加剧隐私泄露,但让最终回答更谨慎
  • 需保护模型内部思考过程,而不仅是输出结果

我们研究了大型推理模型作为个人代理时的隐私泄露问题。与最终输出不同,推理链常被视为内部且安全。我们挑战这一假设,发现推理链频繁包含敏感用户数据,可通过提示注入提取或意外泄露至输出。通过探测和代理评估,我们表明测试时计算方法(尤其是增加推理步数)会放大此类泄露。虽然增加计算预算使模型在最终答案上更谨慎,但也导致其推理更冗长,泄露更多。这揭示了一个核心矛盾:推理提升实用性的同时扩大了隐私攻击面。我们认为,安全措施必须延伸至模型的内部思考,而不仅限于输出。

原文摘要 · Abstract (English)

We study privacy leakage in the reasoning traces of large reasoning models used as personal agents. Unlike final outputs, reasoning traces are often assumed to be internal and safe. We challenge this assumption by showing that reasoning traces frequently contain sensitive user data, which can be extracted via prompt injections or accidentally leak into outputs. Through probing and agentic evaluations, we demonstrate that test-time compute approaches, particularly increased reasoning steps, amplify such leakage. While increasing the budget of those test-time compute approaches makes models more cautious in their final answers, it also leads them to reason more verbosely and leak more in their own thinking. This reveals a core tension: reasoning improves utility but enlarges the privacy attack surface. We argue that safety efforts must extend to the model's internal thinking, not just its outputs.

隐私安全推理链大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。