arXiv:2506.15266cs.CL2025-06EMNLP被引 2

提出首个韩语司法文书脱敏框架,兼顾法律合规与高效处理

Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments

  • 构建首个韩语司法文书标注数据集,明确实体类型
  • 提出系统化个人信息分类体系,适配法律定义
  • 基于深度神经网络实现端到端精准脱敏,性能领先

为平衡司法公开与个人数据保护,韩国司法机构要求公开前对判决文书进行脱敏。然而现有流程难以规模化处理判决文书,且法律中对个人标识信息的定义模糊,不适用于技术实现。为此,我们提出Thunder-DeID脱敏框架,符合相关法律与实践。具体包括:(i) 构建并发布首个含标注判决文书及实体提及列表的韩语法律数据集;(ii) 提出系统化的个人可识别信息(PII)分类体系;(iii) 开发基于深度神经网络(DNN)的端到端脱敏管道。实验表明,该模型在判决文书脱敏任务中达到当前最优性能。

原文摘要 · Abstract (English)

To ensure a balance between open access to justice and personal data protection, the South Korean judiciary mandates the de-identification of court judgments before they can be publicly disclosed. However, the current de-identification process is inadequate for handling court judgments at scale while adhering to strict legal requirements. Additionally, the legal definitions and categorizations of personal identifiers are vague and not well-suited for technical solutions. To tackle these challenges, we propose a de-identification framework called Thunder-DeID, which aligns with relevant laws and practices. Specifically, we (i) construct and release the first Korean legal dataset containing annotated judgments along with corresponding lists of entity mentions, (ii) introduce a systematic categorization of Personally Identifiable Information (PII), and (iii) develop an end-to-end deep neural network (DNN)-based de-identification pipeline. Our experimental results demonstrate that our model achieves state-of-the-art performance in the de-identification of court judgments.

脱敏司法文本PII识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。