提出新理论与生成模型,精准发现高维数据中的异常子空间。
Adversarial Subspace Generation for Outlier Detection in High-Dimensional Data
- 基于新理论将子空间发现转为随机优化问题,避免盲目搜索。
- 在42个真实数据集上显著提升单类分类性能,优于现有方法。
- 适合处理高维数据的异常检测与聚类任务,尤其擅长复杂结构识别。
高维表格数据中的异常检测面临挑战,因数据常分布于多个低维子空间——即多重视角效应(MV)。传统方法依赖启发式搜索,难以准确捕捉数据真实结构。本文提出近视子空间理论(MST),数学化定义MV效应,并将子空间选择建模为随机优化问题。基于此,提出V-GAN生成模型,通过训练求解该问题,无需遍历特征空间即可保留数据内在结构。在42个真实数据集上的实验表明,使用V-GAN子空间构建集成方法可显著提升单类分类性能,优于现有子空间选择、特征选择和嵌入方法。合成数据实验进一步显示,V-GAN在子空间识别精度与扩展性上均优于其他方法。结果验证了理论有效性,并证明其在高维场景下的实用性。
原文摘要 · Abstract (English)
Outlier detection in high-dimensional tabular data is challenging since data is often distributed across multiple lower-dimensional subspaces -- a phenomenon known as the Multiple Views effect (MV). This effect led to a large body of research focused on mining such subspaces, known as subspace selection. However, as the precise nature of the MV effect was not well understood, traditional methods had to rely on heuristic-driven search schemes that struggle to accurately capture the true structure of the data. Properly identifying these subspaces is critical for unsupervised tasks such as outlier detection or clustering, where misrepresenting the underlying data structure can hinder the performance. We introduce Myopic Subspace Theory (MST), a new theoretical framework that mathematically formulates the Multiple Views effect and writes subspace selection as a stochastic optimization problem. Based on MST, we introduce V-GAN, a generative method trained to solve such an optimization problem. This approach avoids any exhaustive search over the feature space while ensuring that the intrinsic data structure is preserved. Experiments on 42 real-world datasets show that using V-GAN subspaces to build ensemble methods leads to a significant increase in one-class classification performance -- compared to existing subspace selection, feature selection, and embedding methods. Further experiments on synthetic data show that V-GAN identifies subspaces more accurately while scaling better than other relevant subspace selection methods. These results confirm the theoretical guarantees of our approach and also highlight its practical viability in high-dimensional settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。