首次构建日文敏感个人信息检测数据集并训练高效识别模型。
Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models

- 用大模型辅助标注构建日文敏感信息数据集
- 训练的分类器能有效识别日文中的敏感信息
- 为日文LLM隐私合规提供关键技术支持
大型语言模型的预训练语料库中可能包含敏感个人数据。为确保符合隐私法规并防止信息泄露,检测和过滤此类信息至关重要。然而,与英语等语言相比,针对日语的敏感个人信息研究仍十分有限。本研究聚焦日本《个人信息保护法》(APPI)定义的特殊保护个人信息(SCPI),利用大模型辅助标注构建了SCPI数据集,并训练机器学习模型以快速检测文本中的SCPI。结果表明,所提出的分类器能有效识别与SCPI相关的信息。本研究是首次探索日语文本中SCPI检测的工作,揭示了准确识别的挑战。
原文摘要 · Abstract (English)
Sensitive personal information can appear in large-scale pre-training corpora for large language models (LLMs). Detecting and filtering such information is therefore essential to ensure compliance with privacy regulations and prevent unintended information leakage. However, in contrast to English and other languages, research into sensitive personal information has been limited in the Japanese language. In this study, we focus on sensitive personal data defined as special care-required personal information (SCPI) under Japan's Act on the Protection of Personal Information (APPI). We construct an SCPI dataset using LLM-based annotation and train machine learning models to rapidly detect SCPI in text. As a result, our SCPI classifier can effectively identify information related to SCPI. This study is the first to explore SCPI detection in Japanese text corpora, highlighting the challenges of accurate detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。