arXiv:2608.21467stat.MLcs.IT2026-08

提出用高斯-埃米特积分法高效计算高斯混合熵,提升强化学习中动作优化精度。

Gauss--Hermite Quadrature for Gaussian-Mixture Entropy with an Action-Space Hermite Surrogate

论文配图:Gauss--Hermite Quadrature for Gaussian-Mixture Entropy with an Action-Space Hermite Surrogate
图 1 · 摘自论文原文
  • 用高斯-埃米特积分逼近高斯混合微分熵,阶数控制精度。
  • 在雷达指向任务中,二阶埃尔米特代理模型误差和优化遗憾显著低于泰勒展开法。
  • 适用于连续动作空间的重复优化场景,如强化学习策略更新。

高斯分布常用于建模信号与状态的不确定性,而高斯混合则适用于多峰分布。与单一高斯不同,高斯混合通常无微分熵的闭式表达,需数值近似。本文提出一种高斯-埃米特积分方法来计算高斯混合微分熵,积分阶数控制近似分辨率。该方法在一维和二维高斯混合基准上,与泰勒近似、解析熵界及数值积分参考结果对比验证。针对连续动作空间的重复优化,我们还提出动作空间的埃尔米特多项式代理模型。在雷达指向任务中,其二阶形式相比基于名义动作局部导数的二阶泰勒代理,虽每重规划步仍仅使用九次目标函数直接评估,但显著降低代理误差与优化遗憾,同时提升指向性能。

原文摘要 · Abstract (English)

Gaussian distributions are used to model uncertainty in signals and states, and Gaussian mixtures are often used when the underlying distribution is multimodal. Unlike a single Gaussian, a Gaussian mixture generally has no closed-form expression for differential entropy and therefore requires numerical approximation. We propose a Gauss--Hermite quadrature method for evaluating Gaussian mixture differential entropy. The quadrature order controls the numerical resolution of the approximation. The method is evaluated on one- and two-dimensional Gaussian mixture benchmarks against Taylor approximations, analytic entropy bounds, and numerical integration references. For repeated optimization over continuous actions, we also propose a Hermite polynomial surrogate in action space. In a radar pointing benchmark, its second-order form achieves substantially lower surrogate error and optimizer regret than a second-order Taylor surrogate based on local derivatives at the nominal action, while both methods use nine direct objective evaluations per replanning step. The Hermite surrogate also improves pointing performance in the tested benchmark.

高斯混合熵计算强化学习代理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。