arXiv:2602.16571cs.CL2026-02被引 1

解决数学辅导对话中数字误判为敏感信息的问题,提升数据可用性。

Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset

  • 用领域感知提示让大模型更准确识别真实敏感信息
  • 数学相关提示使检测准确率从37.9%提升至82.1%
  • 适合教育数据隐私保护与智能教学系统研究者

大规模共享对话数据对教学研究至关重要,但严格的去标识化仍是主要障碍。在数学辅导记录中,数值表达式常与结构化标识符(如日期或编号)相似,导致通用个人身份信息(PII)检测系统过度遮蔽核心教学内容,降低数据效用。本文聚焦“数值模糊”问题,提出首个面向数学辅导对话的PII检测基准数据集MathEd-PII,采用人机协同的大模型标注构建。通过密度分割分析发现,误标集中出现在数学密集区域,证实数值模糊是关键失败模式。比较四种检测策略:Presidio基线与三种基于大模型的方法(基础、数学感知、段落感知提示)。包含数学感知(F1: 0.802)和段落感知提示(F1: 0.821)的方案显著优于基线(F1: 0.379),同时减少数值误报,表明去标识化需融入领域上下文以保留分析价值。本工作提供新基准并证明:教学数据的效用保护型去标识化必须依赖领域感知建模。

原文摘要 · Abstract (English)

Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring transcripts, numeric expressions frequently resemble structured identifiers (e.g., dates or IDs), leading generic Personally Identifiable Information (PII) detection systems to over-redact core instructional content and reduce data utility. This work asks how to detect PII while preserving educational utility, focusing on this "numeric ambiguity" problem. We introduce MathEd-PII, the first benchmark dataset for PII detection in math tutoring dialogues, built with human-in-the-loop LLM annotation. Using density-based segmentation, we show that false PII redactions cluster in math-dense regions, confirming numeric ambiguity as a key failure mode. We then compare four detection strategies: a Presidio baseline and three LLM-based approaches with basic, math-aware, and segment-aware prompting. Domain-aware prompting, including both math-aware (F1: 0.802) and segment-aware versions (F1: 0.821), substantially outperforms the baseline (F1: 0.379) while reducing numeric false positives, demonstrating that de-identification must incorporate domain context to preserve analytic utility. This work provides a new benchmark and evidence that utility-preserving de-identification for tutoring data requires domain-aware modeling.

数据隐私数学教育大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。