跨领域迁移提升敏感信息识别,少量数据也能高效建模。
Cross-Domain Transfer and Few-Shot Learning for Personal Identifiable Information Recognition
- 用多领域数据融合与迁移学习增强模型泛化能力
- 仅需10%训练数据即可在低专业化领域达到高精度
- 法律文本可有效迁移到人物传记,医疗数据难迁移
准确识别个人身份信息(PII)是自动化文本匿名化的关键。本文研究了跨领域模型迁移、多领域数据融合和样本高效学习在PII识别中的有效性。基于医疗(I2B2)、法律(TAB)和传记(Wikipedia)三类标注语料,从本域性能、跨域迁移、数据融合和少样本学习四个维度评估模型表现。结果表明:法律领域数据可有效迁移到传记文本,而医疗领域对跨域迁移具有较强抵抗性;数据融合的效果具有领域特异性;在低专业化领域,仅使用10%的训练数据即可实现高质量识别。
原文摘要 · Abstract (English)
Accurate recognition of personally identifiable information (PII) is central to automated text anonymization. This paper investigates the effectiveness of cross-domain model transfer, multi-domain data fusion, and sample-efficient learning for PII recognition. Using annotated corpora from healthcare (I2B2), legal (TAB), and biography (Wikipedia), we evaluate models across four dimensions: in-domain performance, cross-domain transferability, fusion, and few-shot learning. Results show legal-domain data transfers well to biographical texts, while medical domains resist incoming transfer. Fusion benefits are domain-specific, and high-quality recognition is achievable with only 10% of training data in low-specialization domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。