用大模型分析隐私政策,效果优于传统方法且结果更易懂。
Using LLMs for Automated Privacy Policy Analysis: Prompt Engineering, Fine-Tuning and Explainability
- 结合提示工程与低秩微调提升分类准确率
- 在四个数据集上均显著超越现有最佳方法
- 输出结果可解释性高,适合法律合规等场景
隐私政策广泛用于数字服务,常作为法律要求。已有大量机器学习分类器用于自动识别隐私政策中的关键概念,以辅助生成用户友好摘要或检测合规问题。尽管大语言模型(LLM)在多个自然语言处理任务中表现优异,但其在隐私政策自动化分析中的应用研究仍较少,其潜力尚未被充分探索。为填补这一空白,我们在四个最新的隐私政策语料库和分类体系上,系统评估了基于提示工程和LoRA微调的LLM分类器。实验表明,结合两种方法可显著且一致地超越其他SOTA方法。此外,通过完整性、逻辑性和可理解性三个指标评估可解释性,得分均超过91.1%,表明LLM不仅能提升分类性能,还能增强结果的可解释性。
原文摘要 · Abstract (English)
Privacy policies are widely used by digital services and often required for legal purposes. Many machine learning based classifiers have been developed to automate detection of different concepts in a given privacy policy, which can help facilitate other automated tasks such as producing a more reader-friendly summary and detecting legal compliance issues. Despite the successful applications of large language models (LLMs) to many NLP tasks in various domains, there is very little work studying the use of LLMs for automated privacy policy analysis, therefore, if and how LLMs can help automate privacy policy analysis remains under-explored. To fill this research gap, we conducted a comprehensive evaluation of LLM-based privacy policy concept classifiers, employing both prompt engineering and LoRA (low-rank adaptation) fine-tuning, on four state-of-the-art (SOTA) privacy policy corpora and taxonomies. Our experimental results demonstrated that combining prompt engineering and fine-tuning can make LLM-based classifiers outperform other SOTA methods, \emph{significantly} and \emph{consistently} across privacy policy corpora/taxonomies and concepts. Furthermore, we evaluated the explainability of the LLM-based classifiers using three metrics: completeness, logicality, and comprehensibility. For all three metrics, a score exceeding 91.1\% was observed in our evaluation, indicating that LLMs are not only useful to improve the classification performance, but also to enhance the explainability of detection results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。