用新方法选训练数据,让强化学习的世界模型更准更快。
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
- 基于动作贡献度设计数据筛选指标,避免人为偏见。
- 实测计算效率提升超80%,解决传统方法复杂度难题。
- 适合数据稀缺的强化学习场景,如机器人、自动驾驶。
众多离线与基于模型的强化学习系统依赖世界模型来模拟真实环境。在直接交互成本高、危险或不现实的场景中,世界模型至关重要。其性能与可解释性高度依赖训练数据质量。为此,本文提出一种无偏的数据选择指标——动作沙普利值(Action Shapley),用于高效、公正地筛选训练数据。为降低传统沙普利值计算的指数级复杂度,我们设计了一种随机化动态算法,实验在五个数据受限的真实案例中验证:该算法相较传统方法计算效率提升超过80%。同时,基于动作沙普利值的数据选择策略在性能上持续优于经验性选择方法。
原文摘要 · Abstract (English)
Numerous offline and model-based reinforcement learning systems incorporate world models to emulate the inherent environments. A world model is particularly important in scenarios where direct interactions with the real environment is costly, dangerous, or impractical. The efficacy and interpretability of such world models are notably contingent upon the quality of the underlying training data. In this context, we introduce Action Shapley as an agnostic metric for the judicious and unbiased selection of training data. To facilitate the computation of Action Shapley, we present a randomized dynamic algorithm specifically designed to mitigate the exponential complexity inherent in traditional Shapley value computations. Through empirical validation across five data-constrained real-world case studies, the algorithm demonstrates a computational efficiency improvement exceeding 80\% in comparison to conventional exponential time computations. Furthermore, our Action Shapley-based training data selection policy consistently outperforms ad-hoc training data selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。