解决分布式非参数估计中样本从稀疏到密集的最优率问题
Distributed Nonparametric Estimation: from Sparse to Dense Samples per Terminal
- 设计分层估计协议,利用参数密度估计方法实现最优
- 揭示样本量变化时估计率的相变规律,覆盖全场景
- 适用于密度、回归等多类模型,适合通信受限场景
考虑通信受限下的非参数函数估计问题,每个分布式终端持有多个独立同分布样本。在特定正则条件下,我们刻画了所有情形下的极小极大最优率,并识别出当每终端样本数从稀疏到密集变化时最优率的相变现象。该工作完全解决了此前研究仅限于密集样本或单样本情形的遗留问题。为达到最优率,我们设计了一种分层估计协议,借鉴参数密度估计中的协议。通过信息论方法与强数据处理不等式,结合经典球箱模型,证明了协议的最优性。各类特例(如密度估计、高斯、二值、泊松及异方差回归模型)的最优率可直接得出。
原文摘要 · Abstract (English)
Consider the communication-constrained problem of nonparametric function estimation, in which each distributed terminal holds multiple i.i.d. samples. Under certain regularity assumptions, we characterize the minimax optimal rates for all regimes, and identify phase transitions of the optimal rates as the samples per terminal vary from sparse to dense. This fully solves the problem left open by previous works, whose scopes are limited to regimes with either dense samples or a single sample per terminal. To achieve the optimal rates, we design a layered estimation protocol by exploiting protocols for the parametric density estimation problem. We show the optimality of the protocol using information-theoretic methods and strong data processing inequalities, and incorporating the classic balls and bins model. The optimal rates are immediate for various special cases such as density estimation, Gaussian, binary, Poisson and heteroskedastic regression models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。