提出高效识别分布式学习中异质参数的方法,降低通信开销。
Identifying Heterogeneity in Distributed Learning
- 用重归一化Wald检验和极值对比法检测参数异质性。
- 在异质性稀疏时,极值对比法可处理远超样本量的节点数。
- 方法通信效率高,适合大规模分布式系统中的异常检测。
我们研究了在分布式M估计中以最少数据传输识别异质参数成分的方法。一种基于重归一化Wald检验,只要分布式数据块数K小于最小块样本量且异质性为密集型,就具有一致性。另一种是基于各数据块分量估计参数最大值与最小值之差的极值对比检验(ECT),通过引入样本分割避免M估计带来的偏差累积,在异质性稀疏时,即使K远大于样本量也能保持一致性。ECT操作简便且通信高效。将Wald检验与ECT结合,可在不同异质性稀疏程度下获得更稳健的检验功效。我们进行了大量数值实验,比较了所提方法的族错误率(FWER)和检验力,并通过案例研究验证了方法的实现与有效性。
原文摘要 · Abstract (English)
We study methods for identifying heterogeneous parameter components in distributed M-estimation with minimal data transmission. One is based on a re-normalized Wald test, which is shown to be consistent as long as the number of distributed data blocks $K$ is of a smaller order of the minimum block sample size and the level of heterogeneity is dense. The second one is an extreme contrast test (ECT) based on the difference between the largest and smallest component-wise estimated parameters among data blocks. By introducing a sample splitting procedure, the ECT can avoid the bias accumulation arising from the M-estimation procedures, and exhibits consistency for $K$ being much larger than the sample size while the heterogeneity is sparse. The ECT procedure is easy to operate and communication-efficient. A combination of the Wald and the extreme contrast tests is formulated to attain more robust power under varying levels of sparsity of the heterogeneity. We also conduct intensive numerical experiments to compare the family-wise error rate (FWER) and the power of the proposed methods. Additionally, we conduct a case study to present the implementation and validity of the proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。