arXiv:2604.27550cs.CLcs.AI2026-04ACL

构建首个高精度隐私政策中英对照语料库,助力可读性解读。

APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation

论文配图:APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
图 1 · 摘自论文原文
  • 专家标注139份政策,生成超1.5万条平行文本与三万六千标签。
  • 新框架在可读性与可靠性上超越GPT-4o和LLaMA-3-70B。
  • 适合法律科技、AI可解释性及数据合规研究者使用。

隐私政策对用户理解服务方如何处理个人数据至关重要。然而,这些文件通常冗长复杂,充斥技术术语与法律行话,导致用户无意中接受可能违法的条款。尽管摘要与解读极为关键,但高质量英文平行语料库在法律清晰度与可读性方面仍显不足。为此,我们提出APPSI-139,一个由领域专家精心标注的高质量英语隐私政策语料库,专为摘要与解读任务设计。该语料库包含139份英文隐私政策、15,692条重写后的平行语料,以及覆盖11类数据实践的36,351个细粒度标注标签。同时,我们提出TCSI-pp-V2混合框架,采用交替训练策略,协调多个专家模块,在计算效率与准确率间取得平衡。实验表明,基于APPSI-139语料库与TCSI-pp-V2框架的混合系统,在可读性与可靠性上优于GPT-4o和LLaMA-3-70B等大语言模型。源码与数据集已公开于https://github.com/EnlightenedAI/APPSI-139。

原文摘要 · Abstract (English)

Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and complex, as well as filled with technobabble and legalese, causing users to unknowingly accept terms that may even contradict the law. While summarizing and interpreting these privacy policies is crucial, there is a lack of high-quality English parallel corpus optimized for legal clarity and readability. To address this issue, we introduce APPSI-139, a high-quality English privacy policy corpus meticulously annotated by domain experts, specifically designed for summarization and interpretation tasks. The corpus includes 139 English privacy policies, 15,692 rewritten parallel corpora, and 36,351 fine-grained annotation labels across 11 data practice categories. Concurrently, we propose TCSI-pp-V2, a hybrid privacy policy summarization and interpretation framework that employs an alternating training strategy and coordinates multiple expert modules to effectively balance computational efficiency and accuracy. Experimental results show that the hybrid summarization system built on APPSI-139 corpus and the TCSI-pp-V2 framework outperform large language models, such as GPT-4o and LLaMA-3-70B, in terms of readability and reliability. The source code and dataset are available at https://github.com/EnlightenedAI/APPSI-139.

隐私政策语料库可读性AI解读

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。