arXiv:2411.07317cs.LG2024-11被引 1

用强化学习让生成的临床数据更符合真实研究需求。

SynRL: Aligning Synthetic Clinical Trial Data with Human-preferred Clinical Endpoints Using Reinforcement Learning

  • 用强化学习优化生成器,使数据匹配用户指定的临床终点
  • 在4个真实临床数据集上提升生成数据质量,隐私风险可控
  • 可适配多种生成器,适合需要合规合成数据的研究者

每年数百项临床试验用于评估新疗法,但因隐私和法规限制,患者数据难以共享。为缓解此问题,研究者提出生成合成患者数据的方法。然而,现有方法忽视数据使用需求,未保留临床结局的关键特性,仅依赖生成后的独立评估,无法与生成过程联动。本文提出SynRL,利用强化学习改进患者数据生成器,通过定制化生成满足用户指定的合成数据结果与终点。该方法引入数据价值评判函数评估生成数据质量,并基于评判反馈调整生成器。我们在四个临床试验数据集上验证了该方法的有效性,结果显示其显著提升了生成数据质量,同时保持低隐私风险。此外,SynRL可作为通用框架,适配多种类型的合成数据生成器。代码已公开于https://anonymous.4open.science/r/SynRL-DB0F/。

原文摘要 · Abstract (English)

Each year, hundreds of clinical trials are conducted to evaluate new medical interventions, but sharing patient records from these trials with other institutions can be challenging due to privacy concerns and federal regulations. To help mitigate privacy concerns, researchers have proposed methods for generating synthetic patient data. However, existing approaches for generating synthetic clinical trial data disregard the usage requirements of these data, including maintaining specific properties of clinical outcomes, and only use post hoc assessments that are not coupled with the data generation process. In this paper, we propose SynRL which leverages reinforcement learning to improve the performance of patient data generators by customizing the generated data to meet the user-specified requirements for synthetic data outcomes and endpoints. Our method includes a data value critic function to evaluate the quality of the generated data and uses reinforcement learning to align the data generator with the users' needs based on the critic's feedback. We performed experiments on four clinical trial datasets and demonstrated the advantages of SynRL in improving the quality of the generated synthetic data while keeping the privacy risks low. We also show that SynRL can be utilized as a general framework that can customize data generation of multiple types of synthetic data generators. Our code is available at https://anonymous.4open.science/r/SynRL-DB0F/.

合成数据强化学习临床试验隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。