arXiv:2411.01750cs.LGeess.SP2024-11被引 1

用专家标注心电图数据训练奖励模型,让AI自动设计心脏起搏器。

Show, Don't Tell: Learning Reward Machines from Demonstrations for Reinforcement Learning-Based Cardiac Pacemaker Synthesis

  • 通过神经网络从专家标注的心电图中学习奖励机制
  • 基于学习到的奖励模型,用强化学习生成符合规范的起搏器
  • 适合医疗设备设计与强化学习结合的研究者

心脏起搏器是植入式电子设备,通过电信号调节心跳。随着使用者增加,对带传感器、自适应和长续航功能的需求上升。强化学习(RL)被用于探索起搏器的设计空间、自适应调整及统计验证。正确奖励函数的构建——以奖励机形式表达——是关键环节。2007年波士顿科学公司发布的起搏器规格文档成为多个基于实时自动机与逻辑的形式化描述基础,但这些手动转换难以验证,且要求表达困难。我们提出,心电生理学家更擅长识别患者-起搏器交互中的异常。因此,我们探索从此类标注示范中学习正确性规范,转化为奖励机,并用其训练强化学习代理以合成起搏器。我们利用循环神经网络和变换器架构从标注示范中提取信号作为奖励机。随后,使用该奖励机设计简单起搏器,并基于波士顿科学文档提取的属性进行验证。

原文摘要 · Abstract (English)

An (artificial cardiac) pacemaker is an implantable electronic device that sends electrical impulses to the heart to regulate the heartbeat. As the number of pacemaker users continues to rise, so does the demand for features with additional sensors, adaptability, and improved battery performance. Reinforcement learning (RL) has recently been proposed as a performant algorithm for creative design space exploration, adaptation, and statistical verification of cardiac pacemakers. The design of correct reward functions, expressed as a reward machine, is a key programming activity in this process. In 2007, Boston Scientific published a detailed description of their pacemaker specifications. This document has since formed the basis for several formal characterizations of pacemaker specifications using real-time automata and logic. However, because these translations are done manually, they are challenging to verify. Moreover, capturing requirements in automata or logic is notoriously difficult. We posit that it is significantly easier for domain experts, such as electrophysiologists, to observe and identify abnormalities in electrocardiograms that correspond to patient-pacemaker interactions. Therefore, we explore the possibility of learning correctness specifications from such labeled demonstrations in the form of a reward machine and training an RL agent to synthesize a cardiac pacemaker based on the resulting reward machine. We leverage advances in machine learning to extract signals from labeled demonstrations as reward machines using recurrent neural networks and transformer architectures. These reward machines are then used to design a simple pacemaker with RL. Finally, we validate the resulting pacemaker using properties extracted from the Boston Scientific document.

强化学习起搏器设计奖励机器医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。