arXiv:2409.16621cs.AI2024-09被引 5

用大模型理解隐私政策,让普通人也能看懂数据收集条款。

Entailment-Driven Privacy Policy Classification with LLMs

  • 基于蕴含关系设计标签分类框架,提升理解准确性。
  • 平均F1分数提升11.2%,优于传统方法。
  • 结果可解释,适合普通用户和政策研究者使用。

尽管许多在线服务提供隐私政策供用户阅读以了解个人数据的收集情况,但这些文件通常冗长且复杂。因此,绝大多数用户根本不阅读,导致在不知情的情况下同意数据收集。已有多种尝试通过摘要、自动标注关键段落或提供聊天界面来改善用户体验。随着大语言模型(LLMs)的进展,现在有机会开发更有效的工具来解析隐私政策,帮助用户做出知情决策。本文提出一种基于蕴含关系的LLM框架,将隐私政策段落分类为用户易于理解的标签。实验表明,该框架在平均F1分数上比传统方法提高11.2%,同时提供内在可解释的预测结果。

原文摘要 · Abstract (English)

While many online services provide privacy policies for end users to read and understand what personal data are being collected, these documents are often lengthy and complicated. As a result, the vast majority of users do not read them at all, leading to data collection under uninformed consent. Several attempts have been made to make privacy policies more user friendly by summarising them, providing automatic annotations or labels for key sections, or by offering chat interfaces to ask specific questions. With recent advances in Large Language Models (LLMs), there is an opportunity to develop more effective tools to parse privacy policies and help users make informed decisions. In this paper, we propose an entailment-driven LLM based framework to classify paragraphs of privacy policies into meaningful labels that are easily understood by users. The results demonstrate that our framework outperforms traditional LLM methods, improving the F1 score in average by 11.2%. Additionally, our framework provides inherently explainable and meaningful predictions.

隐私政策大模型自然语言理解可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。