将矩阵补全扩展到分布数据,实现对概率分布的低秩建模与补全。
Low-rank Distributional Matrix Completion

- 用核均值嵌入表示分布,定义分布矩阵的Tucker秩捕捉低秩结构。
- 提出新估计器,在有限样本下实现分布补全,理论误差界可证。
- 适用于高维分布数据补全,如用户行为分布、生物组学数据等场景。
我们研究了一种分布型矩阵补全问题,其中目标矩阵的每个元素是一个概率分布而非标量。在此设定下,仅部分矩阵元素被观测,且对于已观测元素,其底层分布不可直接获取,只能获得从中抽样的有限样本。为表示分布型元素,我们采用核均值嵌入,并引入分布值矩阵的Tucker秩概念以刻画其低秩结构。由于核嵌入的无限维特性带来方法论挑战,我们提出函数型展开算子,将所提出的分布型低秩结构与经典张量的Tucker秩相联系。基于此框架,我们提出一种新的分布型矩阵补全估计器,并建立了非渐近误差界,刻画了估计器的统计性能。在合成数据和一个真实世界应用上的大量实验表明该方法有效。
原文摘要 · Abstract (English)
We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar. In this setting, only a subset of matrix entries is observed, and even for observed entries, the underlying distributions are not directly accessible; instead, we observe finitely many samples drawn from them. To represent distributional entries, we employ kernel mean embeddings and introduce a notion of Tucker rank for distribution-valued matrices to capture their low-rank structure. The infinite-dimensional nature of kernel embeddings poses significant methodological challenges. To address this, we introduce functional unfolding operators that link the proposed distributional low-rank structure to the classical Tucker rank for finite-dimensional tensors. Based on this framework, we propose a novel estimator for distributional matrix completion. We establish non-asymptotic error bounds that characterize the statistical performance of the estimator. Extensive experiments on synthetic data and a real-world application demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。