arXiv:2509.20838cs.CL2025-09EMNLP被引 4

无需训练即可自动隐藏敏感信息,保持文本自然流畅。

Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search

  • 基于树搜索迭代重写句子,逐步处理隐私敏感内容。
  • 在多个数据集上隐私保护效果优于现有方法,同时保持文本自然度。
  • 适合需要快速部署隐私防护的云服务场景。

大型语言模型在云端服务中的广泛应用引发了隐私担忧,用户输入可能无意泄露敏感信息。现有的文本匿名化与去标识化技术(如规则式删除和清理)难以在隐私保护、文本自然性和实用性之间取得平衡。本文提出一种零样本、基于树搜索的迭代句式重写算法,通过结构化搜索与奖励模型引导,系统性地模糊或删除隐私信息,同时保持语义连贯性、相关性和自然性。实验表明,该方法在多个隐私敏感数据集上显著优于现有基线,实现了隐私保护与文本实用性的更优平衡。

原文摘要 · Abstract (English)

The increasing adoption of large language models (LLMs) in cloud-based services has raised significant privacy concerns, as user inputs may inadvertently expose sensitive information. Existing text anonymization and de-identification techniques, such as rule-based redaction and scrubbing, often struggle to balance privacy preservation with text naturalness and utility. In this work, we propose a zero-shot, tree-search-based iterative sentence rewriting algorithm that systematically obfuscates or deletes private information while preserving coherence, relevance, and naturalness. Our method incrementally rewrites privacy-sensitive segments through a structured search guided by a reward model, enabling dynamic exploration of the rewriting space. Experiments on privacy-sensitive datasets show that our approach significantly outperforms existing baselines, achieving a superior balance between privacy protection and utility preservation.

隐私保护文本重写LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。