用大模型实现隐私政策中数据透明度的逐词标注,提升合规评估精度。
Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models
- 构建基于大模型的分步标注流程,融合检索与自纠错机制。
- 在70万份隐私政策上验证,对21项GDPR透明性要求标注准确率显著提升。
- 适合关注数据合规自动化、隐私政策分析的研究者与从业者。
确保个人数据处理的透明性是《通用数据保护条例》(GDPR)的核心要求。然而,由于隐私政策语言复杂多样,大规模合规评估仍具挑战性。人工审计耗时且不一致,现有自动化方法往往缺乏捕捉细微透明度声明所需的粒度。本文提出一种模块化的大语言模型(LLM)管道,实现针对GDPR透明性要求的细粒度词级标注。该方法结合LLM标注、段落级分类、检索增强生成和自纠正机制,可对21项源自GDPR的透明性要求进行可扩展、上下文感知的标注。为支持实证评估,我们构建了一个包含703,791份英文隐私政策的语料库,并基于全面的GDPR对齐标注体系生成200份人工标注的基准样本。我们提出两级评估方法,涵盖段落级分类与片段级标注质量,并对七种主流大模型在两种标注方案(包括广泛使用的OPP-115数据集)上进行了对比分析。结果表明,分解标注任务并集成针对性检索与分类组件,能显著提升标注准确性,尤其在结构清晰的要求上表现更优。本研究提供了新的实证资源与方法基础,推动自动化透明度合规评估的规模化发展。
原文摘要 · Abstract (English)
Ensuring transparency of data practices related to personal information is a core requirement of the General Data Protection Regulation (GDPR). However, large-scale compliance assessment remains challenging due to the complexity and diversity of privacy policy language. Manual audits are labour-intensive and inconsistent, while current automated methods often lack the granularity required to capture nuanced transparency disclosures. In this paper, we present a modular large language model (LLM)-based pipeline for fine-grained word-level annotation of privacy policies with respect to GDPR transparency requirements. Our approach integrates LLM-driven annotation with passage-level classification, retrieval-augmented generation, and a self-correction mechanism to deliver scalable, context-aware annotations across 21 GDPR-derived transparency requirements. To support empirical evaluation, we compile a corpus of 703,791 English-language privacy policies and generate a ground-truth sample of 200 manually annotated policies based on a comprehensive, GDPR-aligned annotation scheme. We propose a two-tiered evaluation methodology capturing both passage-level classification and span-level annotation quality and conduct a comparative analysis of seven state-of-the-art LLMs on two annotation schemes, including the widely used OPP-115 dataset. The results of our evaluation show that decomposing the annotation task and integrating targeted retrieval and classification components significantly improve annotation accuracy, particularly for well-structured requirements. Our work provides new empirical resources and methodological foundations for advancing automated transparency compliance assessment at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。