arXiv:2602.21496cs.AI2026-02

用智能编辑器改写敏感信息,平衡隐私与模型可用性。

Beyond Refusal: Probing the Limits of Agentic Self-Correction for Semantic Sensitive Information

  • 引入代理编辑器,通过迭代修正而非拒绝回答来处理敏感内容。
  • 敏感信息泄露降低34.6%,仅损失9.8%的模型实用性。
  • 大模型靠补充细节提升安全,小模型则靠删减文本,适合不同场景。

尽管结构化个人身份信息(PII)防护已成熟,大型语言模型(LLMs)却带来了新威胁:语义敏感信息(SemSI),即模型推断出敏感身份属性、生成损害声誉的内容或编造错误信息。如何在不破坏模型效用的前提下,让大模型自我调节这些复杂且依赖上下文的敏感信息泄露,仍是未解科学问题。为此,我们提出一种推理时框架SemSIEdit,由一个代理“编辑器”不断批判并重写敏感片段,以保持叙事连贯性而非直接拒绝回答。分析发现存在隐私-效用帕累托前沿:该代理重写机制使三类SemSI泄露减少34.6%,仅造成9.8%的效用损失。同时揭示出规模依赖的安全分化现象:大推理模型(如GPT-5)通过增加细节实现安全,而容量受限模型则回归破坏性截断(删除文本)。最后识别出推理悖论:推理能力虽提升了基础风险(因模型可深入推断敏感信息),但也赋予防御机制执行安全重写的可能。

原文摘要 · Abstract (English)

While defenses for structured PII are mature, Large Language Models (LLMs) pose a new threat: Semantic Sensitive Information (SemSI), where models infer sensitive identity attributes, generate reputation-harmful content, or hallucinate potentially wrong information. The capacity of LLMs to self-regulate these complex, context-dependent sensitive information leaks without destroying utility remains an open scientific question. To address this, we introduce SemSIEdit, an inference-time framework where an agentic "Editor" iteratively critiques and rewrites sensitive spans to preserve narrative flow rather than simply refusing to answer. Our analysis reveals a Privacy-Utility Pareto Frontier, where this agentic rewriting reduces leakage by 34.6% across all three SemSI categories while incurring a marginal utility loss of 9.8%. We also uncover a Scale-Dependent Safety Divergence: large reasoning models (e.g., GPT-5) achieve safety through constructive expansion (adding nuance), whereas capacity-constrained models revert to destructive truncation (deleting text). Finally, we identify a Reasoning Paradox: while inference-time reasoning increases baseline risk by enabling the model to make deeper sensitive inferences, it simultaneously empowers the defense to execute safe rewrites.

大模型安全隐私保护自修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。