arXiv:2511.08073cs.LGstat.ML2025-11AAAI

在线回归中花钱降噪,能有效降低预测误差。

Online Linear Regression with Paid Stochastic Features

  • 通过付费调节特征噪声水平,动态优化预测性能。
  • 已知噪声代价时,误差率最优可达√T;未知时为T²ᐟ³。
  • 适用于需权衡数据质量与成本的实时学习场景。

我们研究一种在线线性回归设置,其中观测到的特征向量受噪声污染,学习者可付费以降低噪声水平。实践中,这可能由于使用更昂贵设备获得更高精度特征,或激励数据提供方释放更少隐私的特征所致。假设特征向量独立同分布于某一固定但未知分布,我们衡量学习者相对于最小化包含预测误差和支付成本组合损失的线性预测器的遗憾。当支付与噪声协方差的映射关系已知时,我们证明在忽略对数因子的情况下,√T 的遗憾率是最优的。当噪声协方差未知时,最优遗憾率变为 T²ᐟ³(同样忽略对数因子)。我们的分析利用矩阵鞅浓度不等式,表明经验损失对所有支付方式和线性预测器都一致收敛至期望值。

原文摘要 · Abstract (English)

We study an online linear regression setting in which the observed feature vectors are corrupted by noise and the learner can pay to reduce the noise level. In practice, this may happen for several reasons: for example, because features can be measured more accurately using more expensive equipment, or because data providers can be incentivized to release less private features. Assuming feature vectors are drawn i.i.d. from a fixed but unknown distribution, we measure the learner's regret against the linear predictor minimizing a notion of loss that combines the prediction error and payment. When the mapping between payments and noise covariance is known, we prove that the rate $\sqrt{T}$ is optimal for regret if logarithmic factors are ignored. When the noise covariance is unknown, we show that the optimal regret rate becomes of order $T^{2/3}$ (ignoring log factors). Our analysis leverages matrix martingale concentration, showing that the empirical loss uniformly converges to the expected one for all payments and linear predictors.

在线学习回归噪声控制后悔率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。