提出约束变量投影新方法,提升结构化模型优化效率
Constrained Variable Projection for Structured Problems
- 将变量投影视为双重优化问题,推导可自动微分的梯度公式
- 在稀疏自编码等任务中,比联合优化快30%以上,数据效率更高
- 适合需要高效求解带凸约束的结构化非线性问题的研究者
变量投影是分离式非线性最小二乘问题的经典技术,通过精确消除线性变量,得到简化后的非线性问题。本文将该框架视为更广泛的双层优化问题的特例,为数据科学模型构建了约束变量投影框架:剩余变量受凸约束,被消除的变量来自下层最小二乘问题。通过将变量投影解释为压缩的双层优化问题,我们推导出与自动微分兼容的精确降维梯度公式,并提出了求解该约束降维问题的条件梯度算法。在标准光滑性和紧致性假设下建立了收敛性保证,并讨论了结构化下层变量的扩展。在稀疏自编码、字典学习、盲去卷积和少样本学习上的数值实验表明,该方法相比自然的联合优化基线,在墙时效率和数据效率上均有提升。
原文摘要 · Abstract (English)
Variable projection is a classical technique for separable nonlinear least-squares problems, in which variables that enter linearly are eliminated exactly, yielding a reduced nonlinear problem. By expressing this framework as a particular instance of a broader class of bilevel optimization problems, we develop a constrained variable-projection framework for data-science models, where the remaining variables are subject to convex constraints and the eliminated variables arise from a lower-level least-squares problem. In particular, by interpreting variable projection as a collapsed bilevel optimization problem, we derive exact reduced-gradient formulas compatible with automatic differentiation and propose a conditional-gradient algorithm for the resulting constrained reduced problem. We establish convergence guarantees under standard smoothness and compactness assumptions, and discuss extensions to structured lower-level variables. Numerical experiments on sparse autoencoding, dictionary learning, blind deconvolution, and few-shot learning suggest that the method can improve wall-clock efficiency and data efficiency relative to natural joint-optimization baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。