arXiv:2409.03707cs.CLcs.AI2024-09被引 1
基于BERT区分词重要性,实现细粒度文本隐私保护
A Different Level Text Protection Mechanism With Differential Privacy
- 用BERT识别文本中不同重要程度的词汇
- 在保持高重要词隐私的前提下提升整体文本可用性
- 适合需要精细控制敏感信息泄露的长文本场景
本文提出一种基于BERT预训练模型的文本重要性分级方法,能够识别出文本中不同重要程度的词汇,并验证了该方法的有效性。研究还探讨了对不同重要程度词汇施加相同扰动对整体文本可用性的影响。该方法可应用于长文本的隐私保护,能够在保护敏感信息的同时维持文本的语义连贯性和实用性。
原文摘要 · Abstract (English)
The article introduces a method for extracting words of different degrees of importance based on the BERT pre-training model and proves the effectiveness of this method. The article also discusses the impact of maintaining the same perturbation results for words of different importance on the overall text utility. This method can be applied to long text protection.
文本隐私BERT差分隐私
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。