arXiv:2603.20208cs.CLcs.AI2026-03被引 1

测试AI能否按安全政策精准删密,同时保留原文意思。

RedacBench: Can AI Erase Your Secrets?

  • 构建跨领域红印评估基准,含514篇人工文本与187条安全政策
  • 用8053个标注命题衡量删除敏感信息与保留非敏感信息的效果
  • 揭示当前模型在保真度上仍难达标,适合数据安全研究者使用

现代语言模型能轻易从非结构化文本中提取敏感信息,因此红印(即选择性删除)对数据安全至关重要。然而,现有红印评估基准多聚焦预定义数据类别(如个人身份信息)或特定技术(如掩码)。为此,我们提出RedacBench,一个面向策略条件红印的综合性评估基准,涵盖个人、企业与政府来源的514篇人工撰写文本,配以187条安全政策,评估模型在移除违反政策信息的同时保持原文语义的能力。通过8,053个标注命题量化每篇文本中可推断出的所有信息,实现对安全性(删除敏感命题)与可用性(保留非敏感命题)的双重评估。实验表明,尽管更先进的模型提升了安全性,但保持语义完整性仍是挑战。为促进后续研究,我们发布RedacBench及网页版交互式工具,支持数据定制与评估,地址:https://hyunjunian.github.io/redaction-playground/

原文摘要 · Abstract (English)

Modern language models can readily extract sensitive information from unstructured text, making redaction -- the selective removal of such information -- critical for data security. However, existing benchmarks for redaction typically focus on predefined categories of data such as personally identifiable information (PII) or evaluate specific techniques like masking. To address this limitation, we introduce RedacBench, a comprehensive benchmark for evaluating policy-conditioned redaction across domains and strategies. Constructed from 514 human-authored texts spanning individual, corporate, and government sources, paired with 187 security policies, RedacBench measures a model's ability to selectively remove policy-violating information while preserving the original semantics. We quantify performance using 8,053 annotated propositions that capture all inferable information in each text. This enables assessment of both security -- the removal of sensitive propositions -- and utility -- the preservation of non-sensitive propositions. Experiments across multiple redaction strategies and state-of-the-art language models show that while more advanced models can improve security, preserving utility remains a challenge. To facilitate future research, we release RedacBench along with a web-based playground for dataset customization and evaluation. Available at https://hyunjunian.github.io/redaction-playground/.

红印评测数据安全语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。