arXiv:2409.13721cs.CL2024-09被引 13

用AI自动判断数据行为是否违规,专攻隐私合规难题。

LegiLM: A Fine-Tuned Legal Language Model for Data Compliance

  • 基于GDPR罚单数据微调,专精于隐私合规评估。
  • 在自建基准上检测违规准确率显著优于基线模型。
  • 适合法律科技、合规团队快速生成法律依据与改进建议。

确保国际数据保护标准下的隐私与数据安全合规是一项关键但复杂的任务,通常需要大量法律专业知识。本文提出LegiLM,一种专为数据或信息合规咨询设计的新型法律语言模型。LegiLM利用预训练的GDPR罚单数据集,并经过微调,可自动评估特定行为或事件是否违反数据安全与隐私法规。通过整合涵盖全球数据保护法律、精细标注的政策文件及相关隐私政策的专用数据集,该模型优化了应对数据合规挑战的能力。模型融合了先进的法律推理方法与信息检索增强技术,提升了实际法律咨询场景中的准确性与可靠性。在自建基准数据集上的评估表明,LegiLM在识别法规违规、提供合理法律解释及推荐必要合规修改方面表现优异,树立了AI驱动法律合规解决方案的新基准。相关资源已公开于https://github.com/DAOLegalAI/LegiLM。

原文摘要 · Abstract (English)

Ensuring compliance with international data protection standards for privacy and data security is a crucial but complex task, often requiring substantial legal expertise. This paper introduces LegiLM, a novel legal language model specifically tailored for consulting on data or information compliance. LegiLM leverages a pre-trained GDPR Fines dataset and has been fine-tuned to automatically assess whether particular actions or events breach data security and privacy regulations. By incorporating a specialized dataset that includes global data protection laws, meticulously annotated policy documents, and relevant privacy policies, LegiLM is optimized for addressing data compliance challenges. The model integrates advanced legal reasoning methods and information retrieval enhancements to enhance accuracy and reliability in practical legal consulting scenarios. Our evaluation using a custom benchmark dataset demonstrates that LegiLM excels in detecting data regulation breaches, offering sound legal justifications, and recommending necessary compliance modifications, setting a new benchmark for AI-driven legal compliance solutions. Our resources are publicly available at https://github.com/DAOLegalAI/LegiLM

法律AI数据合规GDPR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。