提出高效矩阵补全方法,用侧信息降低样本需求,噪声下仍稳定有效。
Sample-efficient inductive matrix completion with noise and inexact side-information
- 用谱初始化的非凸梯度下降法,结合侧信息提升恢复效率。
- 在噪声环境下,样本复杂度由侧信息维度决定,而非矩阵整体大小。
- 支持不精确侧信息,误差随偏差线性增长,适合小样本场景使用。
归纳式矩阵补全(IMC)通过引入行和列的侧信息,理论上可将恢复问题的有效维度降至侧信息特征维度,而非原始矩阵规模。然而现有理论在噪声场景下未能实现此优势:无噪时有样本高效保证,而噪声下所需样本量与传统矩阵补全相当。本文填补这一空白,分析一种带谱初始化的非凸投影梯度下降算法,证明在侧信息精确时,该算法可实现线性收敛与稳定恢复,样本复杂度仅依赖于有效侧信息维度。关键在于建立了在该样本量下仍成立的局部正则性条件,尽管观测模式与侧信息子空间不匹配。进一步扩展至不精确侧信息情形,证明相同低样本复杂度依然成立,且估计误差随子空间误设程度最优地增长。为此,还提出一个权衡样本效率与鲁棒性的加权插值方法。仿真与MovieLens数据集实验验证了理论结果,展示了在小样本下利用侧信息的实际优势。
原文摘要 · Abstract (English)
Inductive matrix completion (IMC) is a variant of low-rank matrix completion that incorporates row and column side-information. In principle, it can reduce the effective dimension of the recovery problem from the ambient matrix size to the dimension of the side-information features. Existing theory, however, does not fully realize this advantage in the noisy setting: sample-efficient guarantees only apply to noiseless recovery, while noisy guarantees require sample sizes comparable to ordinary matrix completion. This paper closes this gap for noisy IMC. We analyze a nonconvex projected gradient descent algorithm with spectral initialization and prove that, under exact side-information, it achieves linear convergence and stable recovery at a sample complexity governed by the effective side-information dimension rather than the ambient matrix dimension. The key technical ingredient is a local regularity condition for the IMC loss that holds at this reduced sample size, despite the mismatch between the observation pattern and the side-information subspaces. We further extend the analysis to inexact side-information, showing that the same reduced sample complexity is preserved and that the estimation error degrades optimally with the level of subspace misspecification. Motivated by this trade-off, we also propose a penalized interpolation between IMC and ordinary matrix completion that balances sample efficiency against robustness to imperfect side-information. Simulations and experiments on the MovieLens dataset support the theoretical findings and illustrate the practical benefits of exploiting side-information in low-sample regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。