arXiv:2506.21744cs.LGstat.AP2025-06

联邦学习保护隐私,实现分布式心理测量建模

Federated Item Response Models: A Gradient-driven Privacy-preserving Framework for Distributed Psychometric Estimation

  • 各机构本地计算梯度,只共享加密汇总数据
  • 隐私保护下精度接近中心化模型,抗异常数据能力强
  • 适合教育评估、医疗心理等敏感数据场景

项目反应理论(IRT)广泛用于估计受试者潜在能力与题目难度。传统方法需集中所有原始答题数据,引发隐私与治理风险。本文提出联邦项目反应理论(FedIRT),可在不传输个体数据的前提下实现分布式IRT模型校准,兼顾隐私保护与统计效率。进一步构建用户级差分隐私扩展FedIRT-DP:各站点计算学生级梯度,裁剪至固定范数后仅共享加噪和;服务器添加校准高斯噪声并执行最大后验更新。该机制提供可审计的$(\varepsilon,δ)$学生级隐私保障,并通过裁剪边界与噪声尺度实现单一可调的隐私-效用权衡。相同机制增强对极端响应行(如全零/全一)的鲁棒性。仿真结果显示,FedIRT在不进行数据汇聚的情况下达到主流R包中心化估计的精度;FedIRT-DP在更强隐私约束下仍保持相近精度,并表现出更优的抗污染能力。真实考试数据实证验证了其可行性与一致的题项及站点效应估计。为推动应用,我们开源了R包FedIRT,支持两参数逻辑斯蒂模型(2PL)与部分信用模型(PCM)的联邦与差分隐私训练。

原文摘要 · Abstract (English)

Item Response Theory (IRT) models are widely used to estimate respondents' latent abilities and calibrate item difficulty. Traditional IRT estimation typically requires centralizing all raw responses, raising privacy and governance concerns. We introduce Federated Item Response Theory (FedIRT), a framework that enables distributed calibration of standard IRT models without transferring individual-level data, thereby preserving confidentiality while retaining statistical efficiency. To provide formal protection, we further develop FedIRT-DP, a user-level differentially private extension. Each site computes per-student gradients, clips them to a fixed norm, and shares only masked sums; the server adds calibrated Gaussian noise and performs MAP updates. This yields an auditable $(\varepsilon,δ)$ guarantee at the student level and a single, tunable privacy-utility trade-off via the clipping bound and noise scale. The same mechanism improves robustness to extreme response rows (e.g., all-zeros/ones). Across simulations, FedIRT matches the accuracy of centralized estimators from popular $\texttt{R}$ packages while avoiding data pooling; FedIRT-DP achieves comparable accuracy under stronger privacy and exhibits superior robustness to contamination. An empirical study on a real exam dataset demonstrates practical viability and consistent item and site-effect estimates. To facilitate adoption, we release an open-source $\texttt{R}$ package, $\texttt{FedIRT}$, implementing the two-parameter logistic (2PL) and partial credit models (PCM) with federated and differentially private training.

联邦学习心理测量差分隐私2PL模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。