arXiv:2503.00174cs.LGstat.ML2025-03ICML被引 3

解决生物数据中缺失值非随机的矩阵补全问题,利用源数据提升补全精度。

Optimal Transfer Learning for Missing Not-at-Random Matrix Completion

  • 用带噪声的源矩阵指导目标矩阵的采样,主动选择最有效行和列。
  • 在主动采样下达到理论最优误差率,无需传统依赖的非相干性假设。
  • 适用于生物数据等高维缺失场景,适合有先验数据的研究者。

我们研究了在缺失非随机(MNAR)设置下的矩阵补全迁移学习,该问题源于生物学应用。目标矩阵 $Q$ 存在整行整列缺失,无辅助信息则无法估计。为此,我们使用一个存在噪声且不完整的源矩阵 $P$,其与 $Q$ 在潜在空间中通过特征偏移相关联。考虑行和列的主动与被动采样两种情形。我们在每种情况下建立了逐项估计误差的极小极大下界。所提出的计算高效估计框架在主动采样情形下达到了该下界,利用源数据选择 $Q$ 中最具信息量的行和列进行查询。这避免了被动采样情形下实现速率最优所需的非相干性假设。我们在真实生物数据集上与现有算法对比,验证了方法的有效性。

原文摘要 · Abstract (English)

We study transfer learning for matrix completion in a Missing Not-at-Random (MNAR) setting that is motivated by biological problems. The target matrix $Q$ has entire rows and columns missing, making estimation impossible without side information. To address this, we use a noisy and incomplete source matrix $P$, which relates to $Q$ via a feature shift in latent space. We consider both the active and passive sampling of rows and columns. We establish minimax lower bounds for entrywise estimation error in each setting. Our computationally efficient estimation framework achieves this lower bound for the active setting, which leverages the source data to query the most informative rows and columns of $Q$. This avoids the need for incoherence assumptions required for rate optimality in the passive sampling setting. We demonstrate the effectiveness of our approach through comparisons with existing algorithms on real-world biological datasets.

矩阵补全迁移学习生物数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。