用验证机制提升目标条件决策的可控性,让模型更准地达成指定目标回报。
Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making
- 融合掩码轨迹重建与价值验证,双目标训练共享模型
- 在低覆盖率区域仍能实现目标回报对齐,误差受覆盖范围和验证精度限制
- 适合需要灵活调节保守或激进策略的临床决策等场景
目标条件序列模型为离线决策提供了简单控制接口,但当目标回报位于数据集覆盖不足区域时,该信号可能不可靠。本文提出 Doctor 框架,结合混合序列建模与强化验证机制,实现可控的目标条件离线决策。Doctor 使用一个共享的掩码轨迹 Transformer,同时优化两个互补目标:掩码轨迹重建用于生成候选动作,样本内值学习用于动作价值验证。推理阶段,模型并行采样多个邻近目标回报,生成候选动作,并选择经验证后最接近目标回报的动作。我们分析了该验证引导的选择规则,证明其价值对齐误差受候选值覆盖范围及验证器准确率约束。在 D4RL 与 EpiCare 数据集上的实验表明,Doctor 在高回报覆盖率降低的情况下仍显著提升目标回报对齐能力,标准离线最大化回报任务上保持竞争力,并在模拟临床决策任务中使单一策略可灵活调节保守与激进操作点。结果表明,强化验证能有效提升目标条件策略的可控性。
原文摘要 · Abstract (English)
Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especially when the target return lies in underrepresented regions of the dataset. This paper proposes Doctor, a hybrid sequence modeling and reinforced verification framework for controllable target-conditioned offline decision making. Doctor trains a shared masked trajectory Transformer with two complementary objectives: masked trajectory reconstruction for candidate generation and in-sample value learning for action-value verification. At inference time, the model samples multiple nearby target returns, generates candidate actions in parallel, and selects the action whose verified value is closest to the requested target return. We analyze this verifier-guided selection rule and show that its value-level alignment error is bounded by candidate-value coverage around the target return and verifier accuracy. Experiments on D4RL and EpiCare show that Doctor improves target-return alignment under reduced high-return coverage, remains competitive on standard offline return-maximization benchmarks, and enables a single policy to modulate between conservative and aggressive operating points in a simulated clinical decision-making task. These results suggest that reinforced verification can improve the controllability of target-conditioned policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。