arXiv:2510.22833cs.AIcs.LG2025-10被引 4

让智能体学会评估计算成本,从而更省力地完成任务。

Toward Agents That Reason About Their Computation

  • 训练时让智能体感知自身计算开销并自主控制计算时机。
  • 在相同训练资源下,75%的游戏表现更优,平均计算量减少三分之二。
  • 适合关注节能、高效推理或可解释性强化学习的研究者。

尽管强化学习智能体在诸多复杂任务中可达到超人水平的表现,但其计算效率通常不会随能力提升而改善。与之相反,人类在熟练掌握任务后会逐渐减少认知负担。若智能体能像人一样在学习过程中思考自身的计算开销,是否也能降低计算消耗?本文通过在雅达利学习环境(Arcade Learning Environment)上实验,让智能体感知计算成本并自主决定何时使用算力。结果显示,在相同的训练计算预算下,能自省计算的智能体在75%的游戏上表现更优,且平均计算量减少至原来的三分之一。我们进一步分析具体游戏,揭示了效率提升的关键机制。

原文摘要 · Abstract (English)

While reinforcement learning agents can achieve superhuman performance in many complex tasks, they typically do not become more computationally efficient as they improve. In contrast, humans gradually require less cognitive effort as they become more proficient at a task. If agents could reason about their compute as they learn, could they similarly reduce their computation footprint? If they could, we could have more energy efficient agents or free up compute cycles for other processes like planning. In this paper, we experiment with showing agents the cost of their computation and giving them the ability to control when they use compute. We conduct our experiments on the Arcade Learning Environment, and our results demonstrate that with the same training compute budget, agents that reason about their compute perform better on 75% of games. Furthermore, these agents use three times less compute on average. We analyze individual games and show where agents gain these efficiencies.

强化学习计算效率智能体节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。