用压缩方法解决高维多输出回归的可解释性与计算难题
Solving Sparse \& High-Dimensional-Output Regression via Compression
- 通过引入稀疏性约束,构建可解释的高维输出回归模型
- 提出两阶段优化框架,压缩后仍保持原精度,计算效率显著提升
- 适合需要高效处理高维科学数据的科研与工业场景
多输出回归(MOR)广泛应用于科学数据分析以支持决策。与传统回归不同,MOR需同时预测多个实值输出。然而,输出维度的增加给模型可解释性和计算可扩展性带来挑战。本文提出一种稀疏且高维输出回归(SHORE)模型,通过引入额外的稀疏性约束来提升输出可解释性,并设计了一种基于输出压缩的两阶段优化框架,可高效求解SHORE并具备可证明的精度。理论上,该框架在任意或较弱样本条件下,计算上具有可扩展性,且训练损失与预测损失在压缩前后保持同阶。实验结果进一步验证了理论发现,展示了该框架在效率与准确性上的优势。
原文摘要 · Abstract (English)
Multi-Output Regression (MOR) has been widely used in scientific data analysis for decision-making. Unlike traditional regression models, MOR aims to simultaneously predict multiple real-valued outputs given an input. However, the increasing dimensionality of the outputs poses significant challenges regarding interpretability and computational scalability for modern MOR applications. As a first step to address these challenges, this paper proposes a Sparse \& High-dimensional-Output REgression (SHORE) model by incorporating additional sparsity requirements to resolve the output interpretability, and then designs a computationally efficient two-stage optimization framework capable of solving SHORE with provable accuracy via compression on outputs. Theoretically, we show that the proposed framework is computationally scalable while maintaining the same order of training loss and prediction loss before-and-after compression under arbitrary or relatively weak sample set conditions. Empirically, numerical results further validate the theoretical findings, showcasing the efficiency and accuracy of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。