arXiv:2601.05076cs.AI2026-01被引 5

让大模型推理不泄露隐私,用提示词或微调减少敏感信息暴露

Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models

  • 通过提示词或微调引导模型生成无隐私泄露的推理链
  • 在多种场景下显著降低敏感信息暴露,性能损失小
  • 提出新评测基准,适合构建安全推理系统的研发者使用

大型推理模型(LRMs)通过生成显式的思维链(CoT)提升性能、可靠性和可解释性,但这种透明性带来了严重隐私风险:中间推理过程常泄露个人身份信息(PII),即使最终答案已去敏。本文研究如何通过可部署的干预手段,实现以隐私为先的推理,而非事后删减。我们提出了PII-CoT-Bench,一个带有隐私意识标注的监督数据集,以及一个覆盖真实与对抗性泄露场景的类别平衡评估基准。结果揭示出能力依赖趋势:顶尖模型通过提示词控制受益最大,而较弱模型需微调才能有效降低泄露。两种方法在各类模型和类别中均显著减少PII暴露,且对任务性能影响极小,证明了在不牺牲性能的前提下实现私密推理是可行的。本工作为构建隐私保护型推理系统提供了实用指导。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) improve performance, reliability, and interpretability by generating explicit chain-of-thought (CoT) reasoning, but this transparency introduces a serious privacy risk: intermediate reasoning often leaks personally identifiable information (PII) even when final answers are sanitized. We study how to induce privacy-first reasoning, where models reason without exposing sensitive information, using deployable interventions rather than post-hoc redaction. We introduce PII-CoT-Bench, a supervised dataset with privacy-aware CoT annotations, and a category-balanced evaluation benchmark covering realistic and adversarial leakage scenarios. Our results reveal a capability-dependent trend: state-of-the-art models benefit most from prompt-based controls, whereas weaker models require fine-tuning to achieve meaningful leakage reduction. Across models and categories, both approaches substantially reduce PII exposure with minimal degradation in utility, demonstrating that private reasoning can be achieved without sacrificing performance. Overall, we show that private CoT reasoning can be achieved with minimal utility loss, providing practical guidance for building privacy-preserving reasoning systems.

隐私保护推理链大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。