提出三种新方法,让连续序列分类误差指数级下降。
Exponentially Consistent Statistical Classification of Continuous Sequences with Distribution Uncertainty
- 不依赖具体分布,设计三种测试方案
- 三类测试的错误概率均指数级衰减
- 适用于分布有偏差的真实场景
在多分类问题中,目标是判断测试序列是否来自与M个训练序列之一相同的分布。不同于以往研究通常关注离散序列且假设分布完全匹配的情况,本文研究连续序列在分布不确定下的多分类问题,即测试序列与训练序列的生成分布即使在真实假设下也存在偏差。我们提出了无分布假设的检验方法,并证明了三种不同测试设计(固定长度、顺序、两阶段)下,检验的错误概率均呈指数级快速衰减。首先考虑无零假设的简单情形:测试序列来自与某个训练序列生成分布相近的分布;随后将结果推广至更一般的情形,允许测试序列来自与所有训练序列生成分布都显著不同的分布。
原文摘要 · Abstract (English)
In multiple classification, one aims to determine whether a testing sequence is generated from the same distribution as one of the M training sequences or not. Unlike most of existing studies that focus on discrete-valued sequences with perfect distribution match, we study multiple classification for continuous sequences with distribution uncertainty, where the generating distributions of the testing and training sequences deviate even under the true hypothesis. In particular, we propose distribution free tests and prove that the error probabilities of our tests decay exponentially fast for three different test designs: fixed-length, sequential, and two-phase tests. We first consider the simple case without the null hypothesis, where the testing sequence is known to be generated from a distribution close to the generating distribution of one of the training sequences. Subsequently, we generalize our results to a more general case with the null hypothesis by allowing the testing sequence to be generated from a distribution that is vastly different from the generating distributions of all training sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。