构建无泄露数据集,首次公正评估文档敏感性分类模型性能
Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

- 用16000份外交电报构建去泄露数据集,消除文本残留标记干扰
- BERT在干净数据上准确率达89.14%,远超传统方法的可靠性
- 适合关注安全与合规的AI研究者及企业文档自动化团队
组织文档的自动敏感性分类是关键但未充分解决的问题,误分类可能导致监管违规或安全漏洞。尽管基于AI的方法可替代人工审核,其可靠性取决于训练数据的完整性。该领域普遍存在却未被广泛报告的标签泄露问题:文档内容中残留的分类标记使模型依赖表面捷径而非真正学习内容敏感信号,导致性能评估虚高且不可靠。本文引入战略16K(Strategic 16K),一个从维基解密美国外交档案公共图书馆(PlusD)精心构建的、受控于泄露的16,000条外交电报语料库,并系统性地评估六种模型架构,涵盖经典机器学习与Transformer模型。提出并执行了扩展的泄露移除协议,识别并清除文档中三类残留分类标记。在清洗后的基准测试中,BERT表现最优(准确率89.14%,F1值89.33%),其次为ELECTRA(准确率88.57%,F1值88.90%)。经典模型中,基于TF-IDF与逻辑回归的方法以显著更低的计算成本取得最佳效果。这是首个基于维基解密PlusD、在明确控制泄露条件下构建的可完全复现的敏感性分类基准。
原文摘要 · Abstract (English)
Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data. A pervasive but underreported problem in this domain is label leakage: residual classification markers embedded within document bodies that allow models to exploit surface shortcuts rather than learning genuine content-based sensitivity signals, producing performance estimates that are inflated and unreliable. This paper addresses this problem by introducing Strategic 16K, a carefully constructed, leakage-controlled corpus of 16,000 diplomatic cables sourced from the WikiLeaks Public Library of US Diplomacy (PlusD), and presents a systematic benchmark evaluating six model architectures spanning classical machine learning and transformer-based approaches. We document an extended leakage removal protocol that identifies and eliminates three categories of residual classification markers embedded within document bodies. On the clean benchmark, BERT achieves the strongest performance (Accuracy = 89.14%, F1 = 89.33%), followed by ELECTRA (Accuracy = 88.57%, F1 = 88.90%). Among classical models, TF-IDF with Logistic Regression achieves the strongest performance at significantly lower computational cost. These results constitute the first fully reproducible sensitivity classification benchmark constructed under explicit leakage-controlled conditions from WikiLeaks PlusD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。