破解隐私政策结构谜题,精准识别数据用途关系
PrivSTRUCT: Untangling Data Purpose Compliance of Privacy Policies in Google Play Store

- 构建编码解码框架,保留政策章节结构线索
- 提取数据用途片段数量超现有工具两倍以上
- 揭示第三方数据披露常被模糊归类,透明度不足
现有研究通常将隐私政策视为扁平文本,忽略其逻辑层级结构,导致自动化方法在关联敏感数据与具体用途时容易混淆。为此,我们提出PrivSTRUCT——一种结合编码器与解码器的系统性框架,旨在理清复杂的隐私披露内容。基准测试显示,相较于最先进的PoliGrapher工具,PrivSTRUCT在保持开发者定义结构线索的同时,提取的数据项与用途片段数量超过两倍。通过对3,756个Android应用的大规模数据分析发现:当开发者使用全局定义的目的而非本地细化说明时,其第一方数据收集的用途夸大概率高出20.4%,第三方共享则高出9.7%。更令人担忧的是,如用于分析的金融数据等敏感第三方数据流,常被稀释并错误归入通用或无关类别,暴露出当前目的披露体系的持续缺陷。
原文摘要 · Abstract (English)
Existing research typically treats privacy policies as flat, uniform text, extracting information without regard for the document's logical hierarchy. Disregard for structural cues of section headings designed to guide the reader, often leads automated methods to entangle distinct data practices, particularly when linking sensitive data items to their specific purposes. To address this, we introduce PrivSTRUCT, a novel and systematic encoder and decoder combined framework that to untangle complex privacy disclosures. Benchmarking against the state-of-the-art tool PoliGrapher reveals that PrivSTRUCT robustly extracts more than x2 the number of data item and purpose excerpts while retaining developer-defined structural cues. By applying PrivSTRUCT to a large-scale dataset of 3,756 Android apps, we uncover a critical transparency gap: the probability of developers overstating a data purpose is 20.4% higher for first-party collection and 9.7% higher for third-party sharing when they rely on globally defined purposes rather than specific, locally scoped disclosures. Alarmingly, we find that sensitive third-party data flows such as sharing financial data for analytics are frequently diluted and entangled into generic or unrelated categories, highlighting a persistent failure in the current purpose disclosure landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。