arXiv:2509.01716cs.AIcs.CL2025-09被引 1

用大模型自动解析隐私政策,构建可查可验的隐私信息图谱。

An LLM-enabled semantic-centric framework to consume privacy policies

  • 基于大模型提取隐私政策关键信息,构建带语义标注的知识图谱。
  • 在100个热门网站上构建了可公开使用的隐私图谱,支持合规审计。
  • 生成符合ODRL等标准的正式政策表达,适合安全与法律研究者使用。

当前用户虽声称会阅读服务条款与隐私政策,但因内容复杂而极少实际阅读,数据隐私实践的模糊性阻碍了以用户为中心的网络发展及智能体时代的数据共享。现有研究尝试通过形式化语言与推理验证政策合规性,但缺乏规模化生成或获取形式化政策的方法。本文提出一种以语义为中心的框架,利用先进大语言模型(LLM)自动识别隐私政策中的关键信息,构建基于数据隐私词汇表(DPV)的$/mathit{Pr}^2/mathit{Graph}$知识图谱,用于支持下游任务。同时,基于该流程分析前100大网站,发布了公开资源。我们还展示了如何用$/mathit{Pr}^2/mathit{Graph}$构建符合开放数字权利语言(ODRL)或永久语义数据使用条款(psDToU)的正式政策表示。为评估技术能力,我们邀请法律专家对Policy-IE数据集进行定制标注,并对比不同大模型在本流水线上的表现。结果表明,该方法具备大规模分析在线服务隐私实践的潜力,是审计网络与互联网的重要方向。所有数据集与源代码均已公开,便于复用与改进。

原文摘要 · Abstract (English)

In modern times, people have numerous online accounts, but they rarely read the Terms of Service or Privacy Policy of those sites, despite claiming otherwise, due to the practical difficulty in comprehending them. The mist of data privacy practices forms a major barrier for user-centred Web approaches, and for data sharing and reusing in an agentic world. Existing research proposed methods for using formal languages and reasoning for verifying the compliance of a specified policy, as a potential cure for ignoring privacy policies. However, a critical gap remains in the creation or acquisition of such formal policies at scale. We present a semantic-centric approach for using state-of-the-art large language models (LLM), to automatically identify key information about privacy practices from privacy policies, and construct $\mathit{Pr}^2\mathit{Graph}$, knowledge graph with grounding from Data Privacy Vocabulary (DPV) for privacy practices, to support downstream tasks. Along with the pipeline, the $\mathit{Pr}^2\mathit{Graph}$ for the top-100 popular websites is also released as a public resource, by using the pipeline for analysis. We also demonstrate how the $\mathit{Pr}^2\mathit{Graph}$ can be used to support downstream tasks by constructing formal policy representations such as Open Digital Right Language (ODRL) or perennial semantic Data Terms of Use (psDToU). To evaluate the technology capability, we enriched the Policy-IE dataset by employing legal experts to create custom annotations. We benchmarked the performance of different large language models for our pipeline and verified their capabilities. Overall, they shed light on the possibility of large-scale analysis of online services' privacy practices, as a promising direction to audit the Web and the Internet. We release all datasets and source code as public resources to facilitate reuse and improvement.

隐私保护大模型应用知识图谱数据合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。