评估生成学生数据在辍学支持中的因果隐私可靠性。
Causal-Privacy Audit Workflow for Synthetic and Distilled Data in Dropout Support
- 提出决策导向的因果隐私审计流程,评估多种生成数据。
- DPGNet和蒸馏数据更准确保留经济状况处理效应结构。
- 适合关注教育决策隐私与因果效应对齐的研究者使用。
合成与蒸馏的学生数据被越来越多用于支持隐私保护的学习分析,但在面向决策的辍学支持中其适用性仍不明确。生成数据不仅需保持预测效用或分布相似性,还需保留用于指导咨询、付款计划协助及奖学金决策的经济状况证据。本研究提出CaP-Eval,一种基于固定估计量、时间感知调整设计、估计器集合和实证隐私治理筛查的决策导向因果隐私审计流程。该流程比较原始数据、蒸馏数据、对抗生成、统计生成以及DPGNet隐私导向生成数据在预测效用、处理效应保真度、对不同估计器的鲁棒性及本地训练记录接近度方面的表现。结果显示,与对抗和高斯核基线相比,DPGNet和蒸馏数据更可靠地保留了原始经济状况处理效应结构;其中,epsilon = 10时,非原始IPW和DML偏差最小,而epsilon = 1和epsilon = 5放大了多个经济状况差异。蒸馏数据保持高度忠实但保留最强本地训练记录信号。TabularGNet保留定性方向但存在中等衰减,高斯核压缩效应幅度。结论:预测效用、隐私导向、实证泄露信号与因果保真度存在分歧,生成学生数据在用于决策前需联合审计方向、大小、重叠与发布治理风险。
原文摘要 · Abstract (English)
Synthetic and distilled student data are increasingly used to enable privacy-conscious learning analytics, yet their suitability for decision-facing institutional support remains uncertain. In dropout support, generated data must preserve not only predictive utility or distributional resemblance, but also the financial-status evidence used to guide advising, payment-plan assistance, and scholarship-related decisions. Method: This study introduces CaP-Eval, a decision-facing causal-privacy audit workflow for evaluating generated student data under a fixed estimand, timing-aware adjustment design, estimator set, and empirical privacy-governance screen. The workflow compares original, distilled, adversarial synthetic, statistical synthetic, and DPGNet privacy-oriented generated data on predictive utility, treatment-effect fidelity, robustness to alternative estimators, and local training-record proximity. Results: DPGNet and distilled data preserved the original financial-status treatment-effect structure more reliably than the adversarial and Gaussian Copula baselines. DPGNet preserved full direction and rank agreement across epsilon levels; epsilon = 10 produced the smallest non-original IPW and DML deviations, while epsilon = 1 and epsilon = 5 amplified several financial-status contrasts. Distilled data remained highly faithful but retained the strongest local training-record proximity signal. TabularGNet preserved qualitative directions with moderate attenuation, and Gaussian Copula compressed effect magnitudes. Conclusions: Predictive utility, privacy orientation, empirical disclosure signals, and causal fidelity diverged; generated student data require joint audits of direction, magnitude, overlap, and release-governance risk before decision use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。