arXiv:2509.19242cs.DScs.LG2025-09被引 1

在数据被恶意删改时,仍能准确估计线性回归模型。

Linear Regression under Missing or Corrupted Coordinates

  • 针对坐标级缺失或篡改,提出新信息论下界与高效算法匹配。
  • 即使有更多样本,最优误差仍不为零,且依赖于参数范围。
  • 发现缺失与篡改场景下最优误差相同,知位置无优势。

我们研究高斯协变量下的多变量线性回归,在两种设定中数据可能被删除或被对手篡改,且每坐标最多影响η比例的样本。在不完整数据情形下,对手可查看数据集并删除每个坐标上不超过η比例的样本(强形式的非随机缺失);在被篡改数据情形下,对手任意替换值,且篡改位置对学习者未知。尽管已有大量关于缺失数据的研究,但此类对抗性缺失下的线性回归仍缺乏理论理解,甚至信息论层面亦然。与干净情形不同,此处最优误差不会随样本增加而消失,而是保持为问题参数的正函数。本文主要贡献是:在几乎全部参数范围内,将最优误差刻画至常数因子以内。具体而言,我们建立了新的信息论下界,其与(计算高效的)算法误差相匹配。一个关键推论是:出人意料地,缺失数据与篡改数据情形下的最优误差一致——知道篡改位置并无普遍优势。

原文摘要 · Abstract (English)

We study multivariate linear regression under Gaussian covariates in two settings, where data may be erased or corrupted by an adversary under a coordinate-wise budget. In the incomplete data setting, an adversary may inspect the dataset and delete entries in up to an $η$-fraction of samples per coordinate; a strong form of the Missing Not At Random model. In the corrupted data setting, the adversary instead replaces values arbitrarily, and the corruption locations are unknown to the learner. Despite substantial work on missing data, linear regression under such adversarial missingness remains poorly understood, even information-theoretically. Unlike the clean setting, where estimation error vanishes with more samples, here the optimal error remains a positive function of the problem parameters. Our main contribution is to characterize this error up to constant factors across essentially the entire parameter range. Specifically, we establish novel information-theoretic lower bounds on the achievable error that match the error of (computationally efficient) algorithms. A key implication is that, perhaps surprisingly, the optimal error in the missing data setting matches that in the corruption setting-so knowing the corruption locations offers no general advantage.

线性回归对抗缺失信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。