arXiv:2506.04293cs.LGcs.AI2025-06EMNLP被引 10

用大模型自动生成可解释的临床试验预测特征,提升效率与可信度。

AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents

  • 基于大模型自动生成并优化表格特征,无需人工干预。
  • 在有限迭代下达到或超越当前最优方法性能。
  • 适合需要可解释性的医药研发与临床决策场景。

临床试验对推进医学治疗至关重要,但成本高昂且耗时。准确预测临床试验结果可显著降低研发成本并加速药物发现。尽管近期深度学习模型通过利用非结构化数据展现出潜力,但其黑箱特性、缺乏可解释性以及易受标签泄露影响,限制了其在高风险生物医学场景中的应用。本文提出AutoCT,一种结合大语言模型推理能力与经典机器学习可解释性的新框架。AutoCT能自主生成、评估并优化基于公开信息的表格特征,无需人工参与。该方法采用蒙特卡洛树搜索,迭代优化预测性能。实验表明,AutoCT在少量自修正迭代后,临床试验预测任务表现达到或优于现有最先进方法,确立了一种可扩展、可解释且成本可控的新范式。

原文摘要 · Abstract (English)

Clinical trials are critical for advancing medical treatments but remain prohibitively expensive and time-consuming. Accurate prediction of clinical trial outcomes can significantly reduce research and development costs and accelerate drug discovery. While recent deep learning models have shown promise by leveraging unstructured data, their black-box nature, lack of interpretability, and vulnerability to label leakage limit their practical use in high-stakes biomedical contexts. In this work, we propose AutoCT, a novel framework that combines the reasoning capabilities of large language models with the explainability of classical machine learning. AutoCT autonomously generates, evaluates, and refines tabular features based on public information without human input. Our method uses Monte Carlo Tree Search to iteratively optimize predictive performance. Experimental results show that AutoCT performs on par with or better than SOTA methods on clinical trial prediction tasks within only a limited number of self-refinement iterations, establishing a new paradigm for scalable, interpretable, and cost-efficient clinical trial prediction.

临床试验大模型可解释性自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。