针对噪声数据分类难题,提出基于特征空间残差的分析框架FINDER。
FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals
- 将数据视为随机场,在希尔伯特空间中构建随机特征
- 通过KLE分解实现降噪,利用谱分析区分不同类别数据
- 在阿尔茨海默病分期和森林砍伐检测中达到领先效果
噪声数据(低信噪比、样本量小、采集故障等)仍是分类方法的前沿挑战,兼具理论与实践意义。本文提出FINDER,一种适用于通用分类问题的严谨框架,专为噪声数据设计。FINDER将基础随机分析思想融入特征学习与推断阶段,以最优方式处理所有经验数据固有的随机性。首先将经验数据视为潜在随机场的实现(不假设其具体分布),再将其映射至适当的希尔伯特空间;通过Kosambi-Karhunen-Loéve展开(KLE)将这些随机特征分解为可计算的不可约成分,进而通过特征值分解实现噪声数据上的分类:不同类别的数据分布在不同区域,可通过关联算子的谱特性识别。FINDER在多个数据稀缺的科学领域验证,取得突破性成果:(i) 阿尔茨海默病阶段分类,(ii) 遥感森林砍伐检测。最后讨论FINDER预期超越现有方法的场景、失效模式及其他局限。
原文摘要 · Abstract (English)
''Noisy'' datasets (regimes with low signal to noise ratios, small sample sizes, faulty data collection, etc) remain a key research frontier for classification methods with both theoretical and practical implications. We introduce FINDER, a rigorous framework for analyzing generic classification problems, with tailored algorithms for noisy datasets. FINDER incorporates fundamental stochastic analysis ideas into the feature learning and inference stages to optimally account for the randomness inherent to all empirical datasets. We construct ''stochastic features'' by first viewing empirical datasets as realizations from an underlying random field (without assumptions on its exact distribution) and then mapping them to appropriate Hilbert spaces. The Kosambi-Karhunen-Loéve expansion (KLE) breaks these stochastic features into computable irreducible components, which allow classification over noisy datasets via an eigen-decomposition: data from different classes resides in distinct regions, identified by analyzing the spectrum of the associated operators. We validate FINDER on several challenging, data-deficient scientific domains, producing state of the art breakthroughs in: (i) Alzheimer's Disease stage classification, (ii) Remote sensing detection of deforestation. We end with a discussion on when FINDER is expected to outperform existing methods, its failure modes, and other limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。