用范畴论统一多种决策模型,让局部行为自然扩展为全局一致决策。
Universal Decision Learners
- 基于左/右坎扩张构建通用决策学习框架
- 涵盖规划、强化学习、因果干预等多类决策问题
- 适合对决策理论与数学建模感兴趣的学者
许多决策理论——如规划、强化学习、因果干预、在线学习和博弈论均衡——都能将局部信息转化为全局一致的行为。本文提出一种统一的范畴论表述:通用决策学习器(UDL)通过一对普遍构造,将已观察上下文中的部分决策函子扩展到新上下文。左坎扩张表达展开、聚合与候选生成;右坎扩张表达一致性、约束满足与不动点语义。核心观点并非所有决策问题有相同算法,而是许多决策形式都可视为同一类普遍问题:以标准方式扩展局部行为数据,再刻画全局一致的扩展。本文给出抽象的UDL构造,证明其普遍比较性质,定义坎不变的行为等价与最小抽象,并展示贝尔曼方程、规划递归、因果干预、在线遗憾及均衡如何成为特例。附录进一步详述强化学习情形。
原文摘要 · Abstract (English)
Many theories of decision making -- planning, reinforcement learning, causal intervention, online learning, and game-theoretic equilibrium -- turn local information into globally coherent behavior. This paper proposes a common categorical formulation: a Universal Decision Learner (UDL) extends a partially specified decision functor from observed contexts to new contexts by a pair of universal constructions. Left Kan extensions express rollout, aggregation, and candidate generation; right Kan extensions express consistency, constraint satisfaction, and fixed-point semantics. The central claim is not that every decision problem has the same algorithm, but that many decision formalisms instantiate the same universal problem: extend local behavioral data canonically, then characterize the globally coherent extensions. We give the abstract UDL construction, prove its universal comparison property, define Kan-invariant behavioral equivalence and minimal abstractions, and show how Bellman equations, planning recursions, causal interventions, online regret, and equilibria arise as special cases. The supplementary material develops the reinforcement-learning specialization in more detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。