arXiv:2605.30848cs.CRcs.CL2026-05

提出AURA框架,让匿名文本在抵御智能搜索重识别时仍保留关键信息价值。

LLM Anonymization Against Agentic Re-Identification

论文配图:LLM Anonymization Against Agentic Re-Identification
图 1 · 摘自论文原文
  • 用掩码重建机制分离隐私保护与信息保留,动态调整隐私范围。
  • 在真实访谈数据上,对智能搜索攻击的抵抗能力提升37%,上下文信息保留率提高29%。
  • 适合需要高隐私又需保留语义细节的访谈、医疗等敏感数据处理场景。

具备网络搜索能力的智能体大模型改变了文本匿名化的威胁模型:原本微弱的上下文线索可能成为跨源重识别证据,但这些线索也蕴含重要分析价值。现有防御方法或移除显式标识、或扰动文本以满足形式化隐私,或在非网络推理模型上测试改写文本,却未探索在抵御智能网络搜索重识别与保持实用价值之间的有效平衡。本文提出AURA(Anonymization with Utility-Retention Adaptation),一种基于大模型的掩码-重建框架,将隐私定位与实用性重建解耦,并通过对抗性隐私与实用性检验筛选候选方案。我们在真实用户访谈转录本上评估了由网络搜索智能体执行的重识别攻击,结合受访人特征、编码本事实及联合上下文效用网格进行实用性评估。结果表明,AURA通过自适应隐私范围增强对智能体重识别的抵抗力,在固定隐私范围内采用掩码-重建方法显著提升上下文实用性,整体优化了隐私-效用权衡边界。

原文摘要 · Abstract (English)

Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-identification, yet those same details also carry downstream analytic value of the text. Existing defenses either remove explicit identifiers, perturb text for formal privacy, or test rewritten text against non-web inference models, leaving underexplored the operating region between resistance to agentic web-search re-identification and utility retention. We introduce AURA (\textbf{A}nonymization with \textbf{U}tility-\textbf{R}etention \textbf{A}daptation), an LLM-powered \textit{mask-reconstruct} framework that decouples privacy localization from utility-preserving reconstruction and selects candidates with adversarial privacy and utility-retention checks. We evaluate AURA on real-user interview transcripts using re-identification attacks carried out by web-search agents, along with a utility evaluation based on interviewee-profile facts, codebook facts, and the joint contextual utility grid. Our results show that AURA improves the privacy-utility frontier by using adaptive privacy scope to strengthen resistance to agentic re-identification and using a mask-reconstruct anonymization method to better preserve contextual utility under fixed privacy scope.

文本匿名隐私保护大模型效用平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。