构建可预测环境行为的有限状态机,用于优化随机博弈中的收益。
Anticipating Oblivious Opponents in Stochastic Games
- 用有限信息状态机建模环境策略,状态对应对环境行为的信念。
- 引入一致性机制,确保信念状态与真实历史下的信念保持固定距离。
- 适用于需预判对手动作的场景,如手术操作与装配任务建模。
我们提出一种系统化方法,用于在并发随机博弈中预测环境中被忽略的行动与策略,同时最大化奖励函数。核心贡献是构造一个有限的信息状态机,其字母表覆盖环境的动作。机器的每个状态映射到对环境策略的信念状态。我们引入一致性概念,保证该状态机追踪的信念状态始终与基于完整历史获得的精确信念状态保持在固定距离内。本文提供了一种验证一致性及合成此类机器的方法,成功终止后可生成满足条件的状态机。该信息状态机可转化为马尔可夫决策过程(MDP),作为计算最大化奖励函数最优策略的起点。我们在多个基准示例上进行了实验评估,包括白内障手术和家具组装的人类活动数据,结果表明该方法能有效预测环境策略与行为,从而提升奖励表现。
原文摘要 · Abstract (English)
We present an approach for systematically anticipating the actions and policies employed by \emph{oblivious} environments in concurrent stochastic games, while maximizing a reward function. Our main contribution lies in the synthesis of a finite \emph{information state machine} whose alphabet ranges over the actions of the environment. Each state of the automaton is mapped to a belief state about the policy used by the environment. We introduce a notion of consistency that guarantees that the belief states tracked by our automaton stays within a fixed distance of the precise belief state obtained by knowledge of the full history. We provide methods for checking consistency of an automaton and a synthesis approach which upon successful termination yields such a machine. We show how the information state machine yields an MDP that serves as the starting point for computing optimal policies for maximizing a reward function defined over plays. We present an experimental evaluation over benchmark examples including human activity data for tasks such as cataract surgery and furniture assembly, wherein our approach successfully anticipates the policies and actions of the environment in order to maximize the reward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。