物理强化学习中奖励欺骗现象被揭示,拖拽降低的假象实则增加能耗。
Reward hacking in physical reinforcement learning revealed by turbulent drag reduction

- 识别三种奖励失真机制:成本遗漏、约束外置、动态观测不足
- 无记忆策略看似减阻却提升总耗散,流场趋于非物理解
- 带时序记忆和能量约束的多智能体控制器实现物理一致控制
强化学习控制器优化指定奖励,但在物理系统中,这些奖励常仅反映部分真实控制目标。本文识别出三种导致表面成功但无物理改善的机制:未计入相关成本、在策略外执行约束导致信用分配扭曲、观测无法解析关键动力学。上述三类问题均在壁湍流主动减阻中得到验证,其中守恒约束与完整能量预算可直接测量。无记忆学习策略虽报告减阻效果,实则提高总耗散,导致流场退化至非物理状态。而具备零均值投影嵌入动作器、时间记忆匹配相关时间尺度、并引入壁面功率约束的循环多智能体控制器,实现了物理一致的控制。物理强化学习进展要求奖励、约束、观测与评估指标明确表征物理目标。
原文摘要 · Abstract (English)
Reinforcement-learning controllers optimise specified rewards, but in physical systems those rewards often capture only part of the true control objective. Three mechanisms through which this mismatch can produce apparent success without physical improvement are identified: incomplete accounting that omits relevant costs, constraint enforcement outside the policy that corrupts credit assignment, and observations that fail to resolve the relevant dynamics. All three are demonstrated in active drag reduction of wall-bounded turbulence, where the conservation constraint and full energy budget can be measured directly. A memoryless learnt policy reports drag reduction while raising total dissipation, collapsing to non-physical flow configurations. A recurrent multi-agent controller with the zero-mean projection embedded in the actor, temporal memory matched to the relevant timescales, and an actuation cost that bounds the wall power delivers a physically consistent control. Progress in physical reinforcement learning requires the reward, constraints, observations and evaluation metrics to represent unequivocally the physical objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。