arXiv:2603.29497cs.CL2026-03中稿 · the LREC CALD-pseu…

用大模型提炼小模型,高效评估文本隐私敏感度

Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models

  • 从6750亿参数大模型中蒸馏出仅1.5亿参数的小模型
  • 在10个领域文本上保持与人工标注高度一致的隐私判断
  • 适合需要快速部署的隐私检测与去标识化系统

文本数据的精准隐私评估仍是隐私保护自然语言处理中的关键挑战。近期研究显示,大型语言模型(LLMs)可作为可靠的隐私评估者,与人类判断高度一致;但其计算开销大,难以大规模处理敏感数据。本文通过将Mistral Large 3(675B)的隐私评估能力蒸馏到参数量仅1.5亿的轻量级编码器模型中,解决了这一问题。基于涵盖10个不同领域的大规模隐私标注文本数据集,训练出高效分类器,在显著降低计算成本的同时保持与人工标注的高度一致性。我们在人工标注测试数据上验证了该方法,证明其在去标识化系统评估中具有实际应用价值。

原文摘要 · Abstract (English)

Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs) can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems.

隐私评估模型蒸馏LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。