用历史数据训练决策模型,帮药企规划抗癌临床试验。
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
- 基于31.7万条公开数据构建881个决策片段,训练离线策略模型。
- 奖励加权行为克隆效果最佳,适应症F1达46.2%,严格F1达14.2%。
- 适合医药研发决策、AI辅助临床试验设计的研究者参考。
临床开发是在不确定性下的序列决策,制药方需从异构证据中规划实验组合。本文将肿瘤学临床开发建模为离线决策问题,让智能体根据决策日可得信息预测未来六个月的药物研发试验组合。为此,我们构建了时间序列数据集,整合31.7万条异构公开记录(包括试验注册、监管审查、企业申报、使用数据和流行病学数据),形成45个历史项目中的881个离线决策时段。对比四种离线目标:行为克隆、奖励加权行为克隆、学习奖励训练和基于价值的隐式Q学习,与四个前沿大模型代理在未见药物、赞助商、药物类别和时间分层上进行比较。离线训练模型优于非微调基线,尤其在2025年8月后无污染测试集表现更优。其中,奖励加权行为克隆表现最佳,取得46.2%的适应症F1和14.2%的严格F1,远超最佳工具代理的25.0%和2.1%。结果表明,结构化离线学习可有效教会智能体规划临床试验。
原文摘要 · Abstract (English)
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。