arXiv:2409.14879cs.CLcs.CY2024-09被引 11

用提示工程让大模型自动分析隐私政策,无需训练就能准确提取关键信息。

Privacy Policy Analysis through Prompt Engineering for LLMs

  • 通过设计提示模板,让大模型零样本理解隐私条款。
  • 在OPP-115数据集上达到超过0.8的F1分数,效果稳定可靠。
  • 适合关注数据隐私、合规审查的研究者与开发者使用。

隐私政策常因复杂性而难以理解,阻碍透明度和知情同意。传统机器学习方法需大量资源与领域训练数据,且难以适应新变化。本文提出PAPEL(Privacy Policy Analysis through Prompt Engineering for LLMs),利用大语言模型结合提示工程,实现隐私政策的自动化分析。该框架通过零样本、单样本及少样本学习,结合思维链提示,引导模型高效提取、标注并总结政策内容,无需额外训练。我们在两个任务中验证其有效性:(i) 信息标注,(ii) 冲突分析。实验表明,所用LLaMA与ChatGPT模型在标注任务中取得不低于0.8的F1分数(基于OPP-115黄金标准),性能媲美现有方法,同时显著降低训练成本并提升可扩展性。

原文摘要 · Abstract (English)

Privacy policies are often obfuscated by their complexity, which impedes transparency and informed consent. Conventional machine learning approaches for automatically analyzing these policies demand significant resources and substantial domain-specific training, causing adaptability issues. Moreover, they depend on extensive datasets that may require regular maintenance due to changing privacy concerns. In this paper, we propose, apply, and assess PAPEL (Privacy Policy Analysis through Prompt Engineering for LLMs), a framework harnessing the power of Large Language Models (LLMs) through prompt engineering to automate the analysis of privacy policies. PAPEL aims to streamline the extraction, annotation, and summarization of information from these policies, enhancing their accessibility and comprehensibility without requiring additional model training. By integrating zero-shot, one-shot, and few-shot learning approaches and the chain-of-thought prompting in creating predefined prompts and prompt templates, PAPEL guides LLMs to efficiently dissect, interpret, and synthesize the critical aspects of privacy policies into user-friendly summaries. We demonstrate the effectiveness of PAPEL with two applications: (i) annotation and (ii) contradiction analysis. We assess the ability of several LLaMa and GPT models to identify and articulate data handling practices, offering insights comparable to existing automated analysis approaches while reducing training efforts and increasing the adaptability to new analytical needs. The experiments demonstrate that the LLMs PAPEL utilizes (LLaMA and Chat GPT models) achieve robust performance in privacy policy annotation, with F1 scores reaching 0.8 and above (using the OPP-115 gold standard), underscoring the effectiveness of simpler prompts across various advanced language models.

隐私分析提示工程大模型应用自动化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。