对比强化学习算法在序列实验设计中的表现,发现带丢弃或集成的算法泛化能力更强。
Performance Comparisons of Reinforcement Learning Algorithms for Sequential Experimental Design
- 用强化学习训练智能体,选择信息量最大的实验设计
- 使用丢弃或集成方法的算法泛化性能更优
- 适合关注实验效率与鲁棒性的研究者
近期序列实验设计研究致力于构建策略,以高效导航设计空间并最大化期望信息增益。尽管已有工作实现可计算的策略,但针对能在实验统计特性变化下仍保持良好性能的泛化策略研究仍不足。近年来,序列实验引入强化学习,训练智能体在设计空间中选择最具有信息量的实验方案。然而,不同强化学习算法在训练此类智能体时的优劣尚不明确。本文比较了多种强化学习算法在序列实验设计场景下的表现,发现训练算法直接影响智能体性能,且采用丢弃(dropout)或集成(ensemble)方法的特定算法在实证中展现出优异的泛化能力。
原文摘要 · Abstract (English)
Recent developments in sequential experimental design look to construct a policy that can efficiently navigate the design space, in a way that maximises the expected information gain. Whilst there is work on achieving tractable policies for experimental design problems, there is significantly less work on obtaining policies that are able to generalise well - i.e. able to give good performance despite a change in the underlying statistical properties of the experiments. Conducting experiments sequentially has recently brought about the use of reinforcement learning, where an agent is trained to navigate the design space to select the most informative designs for experimentation. However, there is still a lack of understanding about the benefits and drawbacks of using certain reinforcement learning algorithms to train these agents. In our work, we investigate several reinforcement learning algorithms and their efficacy in producing agents that take maximally informative design decisions in sequential experimental design scenarios. We find that agent performance is impacted depending on the algorithm used for training, and that particular algorithms, using dropout or ensemble approaches, empirically showcase attractive generalisation properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。