arXiv:2608.17573stat.MLcs.LG2026-08

提出在线回归中特征预热的稀疏遗憾下界,揭示其性能瓶颈。

Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate

  • 基于历史数据重加权设计矩阵,重构最小范数预测器。
  • 证明三种规则在单稀疏比较器下遗憾至少为Ω(√d),与维度相关。
  • 适用于关注高维在线学习稀疏性建模的研究者。

在高维在线预测中,最优预测器可能仅依赖少数特征,因此遗憾应随稀疏性而非环境维度增长。特征预热通过从历史数据估计特征权重,并对重新缩放的设计矩阵进行最小范数拟合来实现此目标。Warmuth和Amid在COLT 2023提出,是否存在三种此类规则之一具有竞争性在线遗憾保证。本文采用仅依赖过去数据的自然Moore-Penrose协议,对这一开放问题的稀疏对数形式给出否定回答。分析揭示了一个共同障碍:廉价的干扰插值导致重构低估真正有预测力的坐标。通过精确的目标质量恒等式与双符号论证,将该效应转化为截断预测损失。Hadamard构造在所有三种规则下强制产生Ω(min{T,√d})的遗憾,扩展至固定素数幂及规则选择情形。相反,遗憾受数据秩控制,欧几里得归一化三角构造对幂次单变量预热达到该依赖关系,即使在非负二阶段岭正则化下仍成立;配对岭构造亦覆盖全部三种幂次规则。对冻结语言模型激活的探索性诊断显示,干扰插值、目标权重与损失间存在相同关系。多变量与Pearson边界仍待解决。

原文摘要 · Abstract (English)

In high-dimensional online prediction, the best predictor may depend on only a few features, so regret should scale with sparsity rather than the ambient dimension. Feature priming pursues this goal by estimating feature weights from past data and refitting a minimum-norm predictor on the rescaled design. Warmuth and Amid asked at COLT 2023 whether any of three such rules admits a competitive online regret guarantee. Using the natural Moore--Penrose protocol based only on past data, we give a negative answer to the sparse-logarithmic form of this COLT open problem. Our analysis identifies a common obstruction: cheap nuisance interpolation causes the refit to underweight the truly predictive coordinate. An exact target-mass identity and a two-sign argument turn this effect into clipped prediction loss. Hadamard constructions force $Ω(\min\{T,\sqrt{d}\})$ regret for all three rules against a zero-loss one-sparse comparator, with extensions to fixed prime powers and selectors among the rules. Conversely, regret is controlled by data rank, and a Euclidean-normalized triangular construction matches this dependence for powered univariate priming, even under nonnegative second-stage ridge regularization; a paired ridge construction also covers all three powered rules. Exploratory diagnostics on frozen language-model activations exhibit the same relation among nuisance interpolation, target weight, and loss. The exact multivariate and Pearson frontiers remain open.

在线学习稀疏性后悔下界特征预热

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。