arXiv:2509.25729cs.CLcs.AI2025-09EMNLP被引 1

用可控编码生成隐私保护的合成文本,兼顾安全与可用性。

Controlled Generation for Private Synthetic Text

  • 引入实体感知控制码,通过上下文学习或前缀调优实现可控生成。
  • 在法律和临床数据上验证,隐私保护与文本质量达到良好平衡。
  • 适合医疗、法律等敏感领域,需兼顾隐私与生成质量的场景使用。

文本匿名化对于在医疗、社会服务、法律等高风险领域负责任地发展和部署人工智能至关重要。本文提出一种新型隐私保护合成文本生成方法,融合去标识化原则与藏于明处(HIPS)理论。该方法引入实体感知控制码,通过上下文学习(ICL)或前缀调优实现可控生成。ICL 版本确保隐私水平与底层去标识系统一致,前缀调优版本采用定制掩码策略与损失函数,支持可扩展、高质量生成。在法律与临床数据集上的实验表明,该方法在隐私保护与文本效用之间实现了良好平衡,为敏感领域提供了实用有效的合成文本生成方案。

原文摘要 · Abstract (English)

Text anonymization is essential for responsibly developing and deploying AI in high-stakes domains such as healthcare, social services, and law. In this work, we propose a novel methodology for privacy-preserving synthetic text generation that leverages the principles of de-identification and the Hiding In Plain Sight (HIPS) theory. Our approach introduces entity-aware control codes to guide controllable generation using either in-context learning (ICL) or prefix tuning. The ICL variant ensures privacy levels consistent with the underlying de-identification system, while the prefix tuning variant incorporates a custom masking strategy and loss function to support scalable, high-quality generation. Experiments on legal and clinical datasets demonstrate that our method achieves a strong balance between privacy protection and utility, offering a practical and effective solution for synthetic text generation in sensitive domains.

文本生成隐私保护合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。