用强化学习方法自动选最优提示,提升决策模型泛化能力
Prompt Tuning Decision Transformers with Structured and Scalable Bandits
- 设计结构化多臂老虎机,动态筛选最优轨迹提示
- 在高维环境和分布外场景中均超越现有提示调优方法
- 可直接复用预训练模型作为特征提取器,高效适配新任务
提示调优已成为在离线强化学习中适配大型预训练决策变压器(DT)的关键技术,尤其适用于多任务和少样本场景。提示决策变压器(PDT)通过从专家示范数据中均匀采样轨迹提示实现任务泛化,但未考虑提示的信息量。本文提出一种基于老虎机的提示调优方法,在推理时从示范数据中学习构建最优轨迹提示。我们设计了一种在轨迹提示空间中运作的结构化老虎机架构,实现了提示规模的线性而非组合式增长。此外,我们证明预训练的PDT本身可作为强大的特征提取器,支持跨多种环境的高效奖励建模。理论上建立了后悔界,并实证表明该方法在广泛任务、高维环境及分布外场景中持续提升性能,优于现有的提示调优基线。
原文摘要 · Abstract (English)
Prompt tuning has emerged as a key technique for adapting large pre-trained Decision Transformers (DTs) in offline Reinforcement Learning (RL), particularly in multi-task and few-shot settings. The Prompting Decision Transformer (PDT) enables task generalization via trajectory prompts sampled uniformly from expert demonstrations -- without accounting for prompt informativeness. In this work, we propose a bandit-based prompt-tuning method that learns to construct optimal trajectory prompts from demonstration data at inference time. We devise a structured bandit architecture operating in the trajectory prompt space, achieving linear rather than combinatorial scaling with prompt size. Additionally, we show that the pre-trained PDT itself can serve as a powerful feature extractor for the bandit, enabling efficient reward modeling across various environments. We theoretically establish regret bounds and demonstrate empirically that our method consistently enhances performance across a wide range of tasks, high-dimensional environments, and out-of-distribution scenarios, outperforming existing baselines in prompt tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。