arXiv:2603.02680cs.AI2026-03

用归一化奖励优化大模型在高频决策中的策略一致性。

LLMs for High-Frequency Decision-Making: Normalized Action Reward-Guided Consistency Policy Optimization

  • 通过归一化动作奖励实现更稳定的策略学习。
  • 在无人机追击任务中,独立与复合任务表现均优于基线。
  • 适合需要高精度实时决策的场景,如自动驾驶、机器人控制。

尽管大型语言模型(LLMs)是序列决策智能体的核心,但在高频决策任务中仍存在固有局限。现有方法多聚焦于低频、语义差异显著的离散环境决策(如家庭规划),难以应对高频任务中状态信息频繁微调的挑战,且子任务与复合任务间常出现策略不一致。本文提出归一化动作奖励引导的一致性策略优化(NAR-CP)。首先,从候选动作的环境反馈中获取预设密集奖励,经归一化完成奖励塑造,并理论证明归一化不影响最优策略;其次,利用LLM推断子观察下的候选动作并生成联合策略,通过一致性损失确保全局语义策略与子级语义策略精确对齐。在典型高频任务无人机追击上实验表明,本方法在独立任务与复合任务中均取得优异性能,且具备良好未见任务泛化能力。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) form the cornerstone of sequential decision-making agent development, they have inherent limitations in high-frequency decision tasks. Existing research mainly focuses on discrete embodied decision scenarios with low-frequency and significant semantic differences in state space (e.g., household planning). These methods suffer from limited performance in high-frequency decision-making tasks, since high-precision numerical state information in such tasks undergoes frequent updates with minimal fluctuations, and exhibiting policy misalignment between the learned sub-tasks and composite tasks. To address these issues, this paper proposes Normalized Action Reward guided Consistency Policy Optimization (NAR-CP). 1) Our method first acquires predefined dense rewards from environmental feedback of candidate actions via reward functions, then completes reward shaping through normalization, and theoretically verifies action reward normalization does not impair optimal policy. 2) To reduce policy misalignment in composite tasks, we use LLMs to infer sub-observation candidate actions and generate joint policies, with consistency loss ensuring precise alignment between global semantic policies and sub-semantic policies. Experiments on UAV pursuit, a typical high-frequency task, show our method delivers superior performance on independent and composite tasks with excellent generalization to unseen tasks.

大模型决策高频控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。