研究分布输入下的分类学习理论,为医疗筛查等场景提供算法保障。
Statistical Learning Theory for Distributional Classification
- 用核方法将分布嵌入希尔伯特空间,再用SVM分类
- 证明了SVM在高斯核下的收敛速度和一致性
- 提出新噪声假设,适用于医学筛查等实际应用
在两阶段采样设置下,监督学习中输入为概率分布,但学习阶段只能获取样本而非分布本身,这在基于学习的医疗筛查或因果推断中具有重要应用。该问题特别适合核方法:先通过核均值嵌入(KME)将分布或样本映射到希尔伯特空间,再使用定义在嵌入空间上的核函数进行标准核方法(如支持向量机,SVM)。本文针对此类方法进行理论分析,重点研究分布输入下的分类任务。我们建立了新的奥拉克不等式,并推导出一致性和学习率结果。对于使用铰链损失与高斯核的SVM,我们提出了二分类文献中已知噪声假设的一个新变体,在此假设下可获得学习率。此外,部分技术工具(如希尔伯特空间上高斯核的新特征空间)本身也具有独立意义。
原文摘要 · Abstract (English)
In supervised learning with distributional inputs in the two-stage sampling setup, relevant to applications like learning-based medical screening or causal learning, the inputs (which are probability distributions) are not accessible in the learning phase, but only samples thereof. This problem is particularly amenable to kernel-based learning methods, where the distributions or samples are first embedded into a Hilbert space, often using kernel mean embeddings (KMEs), and then a standard kernel method like Support Vector Machines (SVMs) is applied, using a kernel defined on the embedding Hilbert space. In this work, we contribute to the theoretical analysis of this latter approach, with a particular focus on classification with distributional inputs using SVMs. We establish a new oracle inequality and derive consistency and learning rate results. Furthermore, for SVMs using the hinge loss and Gaussian kernels, we formulate a novel variant of an established noise assumption from the binary classification literature, under which we can establish learning rates. Finally, some of our technical tools like a new feature space for Gaussian kernels on Hilbert spaces are of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。