arXiv:2412.12931cs.LGstat.ML2024-12中稿 · AAAI

从乱序数据中恢复多子空间矩阵,解决数据清洗与去匿名化难题。

Multi-Subspace Matrix Recovery from Permuted Data

  • 分四阶段处理:剔除异常、重建子空间、分类异常、无监督恢复乱序向量
  • 理论证明异常分类可靠,实现多子空间矩阵精准恢复
  • 适用于数据清洗、融合与去匿名化场景,优于现有方法

本文研究从乱序数据中恢复多子空间矩阵的问题:给定一个列向量来自多个低维子空间的矩阵,部分列在元素层面被随机置换,目标是还原原始矩阵。该任务在数据清洗、集成和去匿名化中有广泛应用,但因存在多个子空间及向量元素的置换,传统方法如鲁棒主成分分析难以有效处理。为此,本文提出一种新颖的四阶段算法流程:异常识别、子空间重建、异常分类、无监督感知下的乱序向量恢复。特别地,为异常分类步骤提供了理论保证,确保多子空间矩阵恢复的可靠性。在多个基准测试中,本方法与当前最优算法对比表现出更优性能。

原文摘要 · Abstract (English)

This paper aims to recover a multi-subspace matrix from permuted data: given a matrix, in which the columns are drawn from a union of low-dimensional subspaces and some columns are corrupted by permutations on their entries, recover the original matrix. The task has numerous practical applications such as data cleaning, integration, and de-anonymization, but it remains challenging and cannot be well addressed by existing techniques such as robust principal component analysis because of the presence of multiple subspaces and the permutations on the elements of vectors. To solve the challenge, we develop a novel four-stage algorithm pipeline including outlier identification, subspace reconstruction, outlier classification, and unsupervised sensing for permuted vector recovery. Particularly, we provide theoretical guarantees for the outlier classification step, ensuring reliable multi-subspace matrix recovery. Our pipeline is compared with state-of-the-art competitors on multiple benchmarks and shows superior performance.

矩阵恢复多子空间数据清洗去匿名化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。