arXiv:2602.18716cs.ROcs.AI2026-02

通过时间建模提升资源使用与动作的协同决策能力

Temporal Action Representation Learning for Tactical Resource Control and Subsequent Maneuver Generation

  • 基于互信息的对比学习捕捉资源与动作的时间依赖关系
  • 量化表示实现多模式、时序一致的动作生成,提升资源利用率
  • 适用于能源受限或感知受限的实时战术决策场景

自主机器人系统需推理资源控制及其对后续动作的影响,尤其是在能量预算有限或感知受限的情况下。基于学习的控制方法能有效处理复杂动态,将问题建模为统一离散资源使用与连续动作的混合动作空间。然而,现有混合动作空间方法未能充分捕捉资源使用与动作间的因果依赖,也忽略了战术决策的多模态特性,而这在快速变化的场景中至关重要。本文提出 TART 框架,即用于战术资源控制与后续动作生成的时间动作表征学习方法。TART 基于互信息目标的对比学习,捕捉资源-动作交互中的内在时间依赖性。所学表征被量化为离散代码本条目,用于条件化策略,从而捕获重复出现的战术模式,实现多模态且时序连贯的行为。我们在两个资源部署关键的领域评估 TART:(i) 一个迷宫导航任务,其中有限的离散动作预算可提升机动性;(ii) 高保真空战模拟器中,F-16 代理需协调武器与防御系统与飞行动作。在两个领域中,TART 均持续优于混合动作基线,证明其在有限资源下高效利用并生成上下文感知动作的有效性。

原文摘要 · Abstract (English)

Autonomous robotic systems should reason about resource control and its impact on subsequent maneuvers, especially when operating with limited energy budgets or restricted sensing. Learning-based control is effective in handling complex dynamics and represents the problem as a hybrid action space unifying discrete resource usage and continuous maneuvers. However, prior works on hybrid action space have not sufficiently captured the causal dependencies between resource usage and maneuvers. They have also overlooked the multi-modal nature of tactical decisions, both of which are critical in fast-evolving scenarios. In this paper, we propose TART, a Temporal Action Representation learning framework for Tactical resource control and subsequent maneuver generation. TART leverages contrastive learning based on a mutual information objective, designed to capture inherent temporal dependencies in resource-maneuver interactions. These learned representations are quantized into discrete codebook entries that condition the policy, capturing recurring tactical patterns and enabling multi-modal and temporally coherent behaviors. We evaluate TART in two domains where resource deployment is critical: (i) a maze navigation task where a limited budget of discrete actions provides enhanced mobility, and (ii) a high-fidelity air combat simulator in which an F-16 agent operates weapons and defensive systems in coordination with flight maneuvers. Across both domains, TART consistently outperforms hybrid-action baselines, demonstrating its effectiveness in leveraging limited resources and producing context-aware subsequent maneuvers.

强化学习战术决策资源控制时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。