arXiv:2602.20743cs.CL2026-02ACL

让文本匿名化自动适配隐私与可用性需求

Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization

  • 用提示词优化自动生成适合不同需求的匿名指令
  • 在5个数据集上均优于现有方法,性能接近大模型
  • 能发现新匿名策略,适合多场景隐私保护

文本匿名化是一个高度依赖上下文的问题:隐私保护与信息可用性之间的平衡随数据领域、隐私目标和下游应用而异。现有方法依赖静态的手动设计策略,缺乏灵活性且难以跨领域泛化。本文提出自适应文本匿名化,通过任务特定的提示词优化框架,自动为语言模型生成匿名化指令,实现对不同隐私目标、领域和下游使用模式的动态适应。我们构建了一个涵盖五个不同数据集的基准测试,覆盖多样化的领域、隐私约束和效用目标。在所有设置中,该框架始终优于现有基线,计算高效,在开源模型上表现良好,性能可比肩更大规模的闭源模型。此外,该方法还能发现探索隐私-效用权衡前沿的新策略。

原文摘要 · Abstract (English)

Anonymizing textual documents is a highly context-sensitive problem: the appropriate balance between privacy protection and utility preservation varies with the data domain, privacy objectives, and downstream application. However, existing anonymization methods rely on static, manually designed strategies that lack the flexibility to adjust to diverse requirements and often fail to generalize across domains. We introduce adaptive text anonymization, a new task formulation in which anonymization strategies are automatically adapted to specific privacy-utility requirements. We propose a framework for task-specific prompt optimization that automatically constructs anonymization instructions for language models, enabling adaptation to different privacy goals, domains, and downstream usage patterns. To evaluate our approach, we present a benchmark spanning five datasets with diverse domains, privacy constraints, and utility objectives. Across all evaluated settings, our framework consistently achieves a better privacy-utility trade-off than existing baselines, while remaining computationally efficient and effective on open-source language models, with performance comparable to larger closed-source models. Additionally, we show that our method can discover novel anonymization strategies that explore different points along the privacy-utility trade-off frontier.

文本匿名化提示优化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。