用随机采样加速多分类模型训练,计算量降为原来的1%仍保持精度。
SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models
- 基于无偏小批量梯度估计,每轮迭代固定样本量,不随数据规模增长。
- 在多项式概率模型中,仿真误差可忽略,计算时间不足传统方法的1%。
- 适用于多种受限因变量模型,适合大规模数据分析场景。
多分类模型虽能灵活描述替代关系,但当备选方案或观测数增多时计算成本急剧上升。固定每观测仿真实验预算下,模拟最大似然法引入仿真偏差,且每次优化需全样本似然评估。本文提出随机近似无偏模拟得分(SAUSS),基于条件无偏的小批量得分估计,每轮迭代使用固定大小小批量,不受样本量影响。对于多项式概率模型,接受-拒绝采样可实现任意固定接受次数下的精确条件抽样与无偏得分估计。在局部条件下,平均估计量及SAUSS迭代部分和过程的渐近理论,纳入小批量与仿真变异性,支持随机缩放与插值推断。模拟与实际应用显示,SAUSS在小于1%的计算时间内达到与模拟最大似然相当的结果。该方法可扩展至具有条件期望得分表示及精确条件采样的受限因变量模型。
原文摘要 · Abstract (English)
Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose Stochastic Approximation with Unbiased Simulated Scores (SAUSS), an averaged stochastic approximation based on conditionally unbiased mini-batch score estimates. Each iteration uses a fixed mini-batch regardless of sample size. For multinomial probit, accept-reject sampling provides exact conditional draws and unbiased score estimates for any fixed number of accepted draws. Under local conditions, asymptotic theory for the averaged estimator and the partial-sum process of the SAUSS iterates incorporates mini-batch and simulation variability and supports random-scaling and plug-in inference. In simulations and an application, SAUSS gives comparable results in less than 1% of the computation time of simulated maximum likelihood. SAUSS extends to limited dependent variable models with conditional-expectation score representations and exact conditional sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。