arXiv:2505.16931cs.CL2025-05EMNLP被引 12

轻量级框架Pivot通过上下文信息提升对话中敏感信息识别准确率。

PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues

  • 利用对话上下文简化敏感信息检测任务
  • 构建了最大规模真实教学对话数据集QATD-2k
  • 适合教育数据共享与隐私保护研究者使用

个人身份信息(PII)匿名化是一项高风险任务,制约着开放科学的数据共享。尽管近年来PII识别取得显著进展,但实际应用中错误阈值和召回率/精确率权衡仍限制了匿名化流程的采纳。本文提出PIIvot,一种基于数据上下文知识的轻量化匿名化框架,可有效简化敏感信息检测问题。为验证其有效性,我们还发布了目前最大规模的开源真实教学对话数据集QATD-2k,以支持高质量教育对话数据的需求。

原文摘要 · Abstract (English)

Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds and the recall/precision trade-off still limit the uptake of these anonymization pipelines. We present PIIvot, a lighter-weight framework for PII anonymization that leverages knowledge of the data context to simplify the PII detection problem. To demonstrate its effectiveness, we also contribute QATD-2k, the largest open-source real-world tutoring dataset of its kind, to support the demand for quality educational dialogue data.

隐私保护对话数据轻量框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。