用大模型实现多语言隐私政策审计,发现西班牙应用存在中英文政策不一致问题。
Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps

- 用大模型直接分析多语种隐私政策,无需针对每种语言单独训练。
- 跨语言分类准确率稳定在0.91至0.94之间,支持大规模评估。
- 揭示公共部门用西语、商业应用用英语,导致透明度被掩盖。
自动化隐私政策分析可实现数字生态系统的规模化透明度评估,但现有方法主要依赖英语,难以覆盖多语言环境。本文研究大语言模型(LLMs)是否能在不进行语言特化调整的情况下,将隐私政策分析扩展至非英语场景。我们构建了一个涵盖欧盟24种官方语言的评估语料库,基于两个已有的专家标注数据集(OPP-115 和 MAPP)的翻译版本,并通过自动指标和法律专家评审评估翻译质量。实验显示,基于LLM的个人数据收集类别识别器在跨语言任务中表现稳定,宏平均F1得分介于0.91至0.94之间。随后,我们对西班牙Google Play商店中的2,611个Android应用进行了大规模审计,结合多语言隐私政策分析、隐私标签评估与运行时网络流量检测,发现公共部门应用主要提供西班牙语政策,而流行商业应用则多使用英语。结果显示,声明与实际行为之间存在系统性差异,尤其在公共部门应用中更为明显。总体表明,仅限英语的审计会系统性掩盖多语言环境中的透明度缺陷。
原文摘要 · Abstract (English)
Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, yet existing auditing pipelines remain predominantly English-centric. This limits their ability to systematically evaluate multilingual environments, as in the European Union, where many services disclose privacy practices only in local languages. This paper examines whether large language models (LLMs) can extend privacy policy analysis beyond English without requiring language-specific adaptation, thus empowering large-scale auditing in linguistically diverse app ecosystems. We assemble an evaluation corpus spanning all 24 official EU languages from translated versions of two established expert-annotated datasets (OPP-115 and MAPP) and assess translation fidelity through automated metrics and targeted legal-expert review. Our LLM-based classifier for identifying categories of personal data collection achieves stable cross-lingual performance, with macro-F1 scores ranging between 0.91 and 0.94. We then leverage this capability in a large-scale audit of 2,611 Android applications from the Spanish Google Play Store. Combining multilingual privacy policy analysis with the evaluation of corresponding privacy labels and runtime network traffic exposes an important linguistic barrier: public-sector apps predominantly provide privacy policies in Spanish, whereas popular commercial apps mostly provide them in English. We reveal systematic discrepancies between declared and observed practices, especially in public-sector apps. Overall, our results indicate how English-only privacy audits can systematically obfuscate transparency gaps in multilingual environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。