用核方法提升逆优化建模能力,实现高效学习专家决策目标
Scalable Kernel Inverse Optimization
- 将逆优化目标函数扩展到无限维再生核希尔伯特空间
- 通过重构定理将问题转化为有限维凸优化,保持可解性
- 提出顺序选择算法加速训练,适用于大规模决策学习任务
逆优化(IO)是一种从历史数据中学习专家决策者未知目标函数的框架。本文将IO的目标函数假设类扩展至再生核希尔伯特空间(RKHS),使特征表示进入无限维空间。我们证明了特定训练损失下重构定理的变体成立,从而可将原问题重构成有限维凸优化问题。为解决核方法常见的可扩展性问题,提出了序列选择优化(SSO)算法,以高效训练所提出的核逆优化(KIO)模型。最后,通过在MuJoCo基准上的示范学习任务,验证了KIO模型的泛化能力及SSO算法的有效性。
原文摘要 · Abstract (English)
Inverse Optimization (IO) is a framework for learning the unknown objective function of an expert decision-maker from a past dataset. In this paper, we extend the hypothesis class of IO objective functions to a reproducing kernel Hilbert space (RKHS), thereby enhancing feature representation to an infinite-dimensional space. We demonstrate that a variant of the representer theorem holds for a specific training loss, allowing the reformulation of the problem as a finite-dimensional convex optimization program. To address scalability issues commonly associated with kernel methods, we propose the Sequential Selection Optimization (SSO) algorithm to efficiently train the proposed Kernel Inverse Optimization (KIO) model. Finally, we validate the generalization capabilities of the proposed KIO model and the effectiveness of the SSO algorithm through learning-from-demonstration tasks on the MuJoCo benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。