用共享表示加速能源系统智能调控的快速适应
Meta-RL with Shared Representations Enables Fast Adaptation in Energy Systems
- 双层优化+混合演员-评论家架构,提升样本效率
- 共享状态特征提取器使跨任务适应更快,减少过拟合
- 适合需要快速响应变化的智能建筑能源管理场景
元强化学习可解决传统强化学习在多任务和非平稳环境中的适应慢、泛化差问题。本文提出一种新型元强化学习框架,采用双层优化与混合演员-评论家结构,提升样本效率与跨任务适应能力。通过联合优化演员与评论家网络,元学习一个共享的状态特征提取器,实现高效表征学习并抑制对单一任务或主导模式的过拟合。此外,提出外层与内层演员网络间的参数共享机制,减少重复学习,加速任务重访时的适应过程。该方法在覆盖近十年时间与结构变化的真实建筑能源管理系统数据集上验证,配合任务准备方法促进泛化。实验表明,其任务适应能力更强,性能优于传统RL与现有Meta-RL方法。
原文摘要 · Abstract (English)
Meta-Reinforcement Learning addresses the critical limitations of conventional Reinforcement Learning in multi-task and non-stationary environments by enabling fast policy adaptation and improved generalization. We introduce a novel Meta-RL framework that integrates a bi-level optimization scheme with a hybrid actor-critic architecture specially designed to enhance sample efficiency and inter-task adaptability. To improve knowledge transfer, we meta-learn a shared state feature extractor jointly optimized across actor and critic networks, providing efficient representation learning and limiting overfitting to individual tasks or dominant profiles. Additionally, we propose a parameter-sharing mechanism between the outer- and inner-loop actor networks, to reduce redundant learning and accelerate adaptation during task revisitation. The approach is validated on a real-world Building Energy Management Systems dataset covering nearly a decade of temporal and structural variability, for which we propose a task preparation method to promote generalization. Experiments demonstrate effective task adaptation and better performance compared to conventional RL and Meta-RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。