arXiv:2511.19628stat.MLcs.LG2025-11

MCMC在任意目标函数下表现依赖似然尖锐度,调节可提升正则化效果。

Optimization and Regularization Under Arbitrary Objectives

  • 用双块MCMC交替采样,通过似然尖锐度控制正则化强度。
  • 似然曲率决定模型性能和数据推断的正则程度。
  • 适用于强化学习中的导航与井字棋任务,可替代复杂采样流程。

本研究探讨将马尔可夫链蒙特卡洛(MCMC)方法应用于任意目标函数时的局限性,聚焦于交替使用Metropolis-Hastings与Gibbs采样的两区块框架。尽管此类方法常被认为有利于实现数据驱动的正则化,但其性能关键取决于所采用似然形式的尖锐度。通过引入尖锐度参数,并探索与目标函数成比例的替代似然形式,我们表明似然曲率同时影响模型的样本内表现及训练数据所推断出的正则化程度。在强化学习任务中进行实证分析:包括一个导航问题和井字棋游戏。研究进一步对经典黑杰克游戏中的极端似然尖锐度展开独立分析,其中两区块MCMC框架的第一步被迭代优化步骤取代。结果表明,该混合方法性能几乎等同于原始MCMC框架,说明过高的似然尖锐度会迫使后验质量坍缩至单一主导模式。

原文摘要 · Abstract (English)

This study investigates the limitations of applying Markov Chain Monte Carlo (MCMC) methods to arbitrary objective functions, focusing on a two-block MCMC framework which alternates between Metropolis-Hastings and Gibbs sampling. While such approaches are often considered advantageous for enabling data-driven regularization, we show that their performance critically depends on the sharpness of the employed likelihood form. By introducing a sharpness parameter and exploring alternative likelihood formulations proportional to the target objective function, we demonstrate how likelihood curvature governs both in-sample performance and the degree of regularization inferred by the training data. Empirical applications are conducted on reinforcement learning tasks: including a navigation problem and the game of tic-tac-toe. The study concludes with a separate analysis examining the implications of extreme likelihood sharpness on arbitrary objective functions stemming from the classic game of blackjack, where the first block of the two-block MCMC framework is replaced with an iterative optimization step. The resulting hybrid approach achieves performance nearly identical to the original MCMC framework, indicating that excessive likelihood sharpness effectively collapses posterior mass onto a single dominant mode.

MCMC正则化强化学习似然尖锐度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。