arXiv:2605.10018cs.LG2026-05

物理先验能显著减少决策所需数据,尤其在小样本时更有效。

The Value of Mechanistic Priors in Sequential Decision Making

论文配图:The Value of Mechanistic Priors in Sequential Decision Making
图 1 · 摘自论文原文
  • 用占用加权偏差衡量模型与真实最优策略的相似性
  • 大样本下理论证明可降低采样复杂度至原基准的 $H(μ)/H_{\mathrm{mech}}$
  • 小样本时验证了错误先验的代价,适合医疗等安全关键场景

混合机制模型(物理先验+学习残差)有望减少优质决策所需数据量,但缺乏可计算的检验标准。本文在渐近和初始阶段分别分析机制先验的价值。提出机制信息概念——模型推荐策略 $\hatπ$ 与真实最优策略 $π^*$ 的互信息,由占用加权偏差 $B_μ$ 衡量。在大样本($N$ 大)情形,匹配界表明贝叶斯后悔随残差熵 $H_{\mathrm{mech}}$ 增长,相比无先验基线,理论采样复杂度降低 $H(μ)/H_{\mathrm{mech}}$。同时提供模型证书以实证评估采样效率。在小样本($N$ 小)的临床相关烧入阶段,建立错误信心先验的惩罚下界。通过基于已发表 FOLFOX 药代动力学数据的 5-氟尿嘧啶(5-FU)给药模拟,在5种情形中验证了渐近与烧入边界,显示混合先验在烧入阶段带来显著采样效率提升。最后对比大语言模型(LLM)先验,发现其机制信息严重损失,从而支持安全关键应用中仅使用物理基础先验。

原文摘要 · Abstract (English)

Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criterion to test this. We characterize the value of mechanistic priors in sequential decision-making within both asymptotic and burn-in regimes. To formalize this, we introduce the mechanistic information of a model -- the mutual information between the model's recommended policy $\hatπ$ and the true optimal policy $π^*$ -- quantified via an occupancy-weighted bias $B_μ$. In the asymptotic regime (large $N$), matched bounds reveal that Bayesian regret scales with the residual entropy $H_{\mathrm{mech}}$, delivering a theoretical sample complexity reduction of $H(μ)/H_{\mathrm{mech}}$ compared to an uninformed baseline. Furthermore, we provide a model certificate to determine empirical sample efficiency. Complementarily, in the clinically relevant burn-in regime (small $N$), we establish a lower bound on the penalty incurred by confidently wrong priors. We demonstrate both the asymptotic and burn-in bounds across 5-fluorouracil (5-FU) dosing simulations motivated by published FOLFOX pharmacokinetic data, where a hybrid prior yields large sample-efficiency gains in the burn-in regime. Finally, we contrast these grounded models with LLM priors, demonstrating that LLMs can suffer severe losses in mechanistic information, thereby motivating the exclusive use of physically-grounded priors for safety-critical applications.

强化学习机制先验小样本决策医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。