轻量级框架Pivot通过上下文信息提升对话中敏感信息识别准确率。
PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues
- 利用对话上下文简化敏感信息检测任务
- 构建了最大规模真实教学对话数据集QATD-2k
- 适合教育数据共享与隐私保护研究者使用
个人身份信息(PII)匿名化是一项高风险任务,制约着开放科学的数据共享。尽管近年来PII识别取得显著进展,但实际应用中错误阈值和召回率/精确率权衡仍限制了匿名化流程的采纳。本文提出PIIvot,一种基于数据上下文知识的轻量化匿名化框架,可有效简化敏感信息检测问题。为验证其有效性,我们还发布了目前最大规模的开源真实教学对话数据集QATD-2k,以支持高质量教育对话数据的需求。
原文摘要 · Abstract (English)
Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds and the recall/precision trade-off still limit the uptake of these anonymization pipelines. We present PIIvot, a lighter-weight framework for PII anonymization that leverages knowledge of the data context to simplify the PII detection problem. To demonstrate its effectiveness, we also contribute QATD-2k, the largest open-source real-world tutoring dataset of its kind, to support the demand for quality educational dialogue data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。