用费舍尔信息惩罚纠正数据分布变化,提升模型评估可靠性。
Causal Covariate Shift Correction using Fisher information penalty
- 通过费舍尔信息累积训练批次的数据密度知识
- 相比全数据基线,准确率最高提升20.3%(批内)
- 适合处理时序分布变化的机器学习场景
训练数据在不同批次间特征分布随时间演变,导致交叉验证偏差,使模型选择与评估不可靠。本文提出因果协变量偏移校正(C³),从分布式密度估计角度建模时序数据分布变化。利用费舍尔信息累积每个训练批次的数据密度知识,并对后续所有批次的损失施加惩罚。该方法在批量基准下最大提升20.3%准确率,在折叠基准下最低提升5.9%,相较于全数据基线提升12.9%。
原文摘要 · Abstract (English)
Evolving feature densities across batches of training data bias cross-validation, making model selection and assessment unreliable (\cite{sugiyama2012machine}). This work takes a distributed density estimation angle to the training setting where data are temporally distributed. \textit{Causal Covariate Shift Correction ($C^{3}$)}, accumulates knowledge about the data density of a training batch using Fisher Information, and using it to penalize the loss in all subsequent batches. The penalty improves accuracy by $12.9\%$ over the full-dataset baseline, by $20.3\%$ accuracy at maximum in batchwise and $5.9\%$ at minimum in foldwise benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。