arXiv:2607.17075cs.CRcs.CL2026-07

大模型已能替代专业工具,高效分析隐私政策合规性。

A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs

论文配图:A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs
图 1 · 摘自论文原文
  • 用大模型直接处理隐私政策分析任务,无需额外训练
  • 在实体识别上准确率超80%,部分任务优于传统工具
  • 适合法律科技、数据合规从业者快速评估政策

大语言模型的出现显著改变了隐私政策与数据合规分析的研究范式,使此前需依赖专用领域工具的任务得以实现。然而,大模型能否真正复现以往研究中多样化的功能与方法仍不明确。本文首次系统评估现成大模型(GPT-5.2 和 Gemini-2.5)是否可替代专业化隐私分析工具。我们选取六款代表性工具,覆盖三类核心功能:矛盾检测、法规合规分析、隐私政策摘要与聚合,并涵盖三类中间任务:基于元组的结构化数据提取、语义角色标注(SRL)及人工标签生成。在自建的10份隐私政策数据集上,直接提示大模型执行对应任务,评估其在无工程改造或领域训练下是否具备工具级功能。结果表明,大模型在各项功能上均达到或超过传统工具表现:在第一方数据收集主体的人工标注中,平均精确率为81.8%,召回率为70.9%;第三方共享主体标注中,平均精确率为91.4%,召回率为70.8%(基于OPP-115数据集)。总体表明,大模型可有效承担此前需专用工具完成的多样化隐私政策与法规分析任务。

原文摘要 · Abstract (English)

The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously required specialized, domain-specific tools. However, it remains unclear to what extent LLMs can truly replicate the diverse functionalities, and the wide range of methodologies and analysis offered by prior work. In this paper, we conduct the first systematic evaluation of whether off-the-shelf LLMs can replace specialized privacy analysis tools. We study six representative tools spanning three major functionalities: contradiction detection, regulatory compliance analysis, and privacy policy summarization and aggregation, and across three intermediate tasks: structured data extraction using tuples, Semantic Role Labeling (SRL) and manual privacy policy labeling. We compare the performance of two state-of-the-art LLMs (GPT-5.2 and Gemini-2.5 in various configurations) against the tools by directly prompting the models to perform corresponding functionalities and tasks on a custom dataset of 10 privacy policies, allowing us to assess whether off-the-shelf models can produce tool-specific functionalities without further engineering or domain-specific training, major limitations in prior work. Our results show that LLMs consistently match or exceed the capabilities of existing tools across the functionalities. In manual labeling of first-party collection entities, LLMs achieved an average precision of 81.8% and recall of 70.9%, while for labeling of third-party sharing entities, they achieved an average precision of 91.4% and recall of 70.8% compared to the OPP-115 dataset. Overall, our findings indicate that LLMs can effectively perform a broad range of functionalities and tasks in privacy policy and regulation analysis that previously required specialized tools.

大模型隐私政策合规分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。