用大模型自动构建可落地的临床评分系统,兼顾准确与实用。
Automatic Construction of Clinical Scoring Systems with LLM Agents
- 用大模型生成候选规则,再通过数据验证筛选可部署规则
- 在8个任务中表现优于现有方法,性能接近灵活模型
- 适合需要可解释、易记忆临床评分的医生和研究者
现代临床实践依赖基于证据的指南,以简洁的评分系统形式呈现,由少量可解释的决策规则组成。尽管机器学习模型性能强,但因与临床工作流不匹配(如难以记忆、审计困难、无法床边使用),难以实际应用。我们指出,问题不在预测能力不足,而在于优化目标所选模型类别本身不适配指南部署。可部署的指南常为单位权重的临床清单,通过二元规则求和并阈值化形成,但此类评分需在指数级庞大的规则组合空间中搜索,难度极高。本文提出AgentScore,利用大模型生成候选规则,并通过确定性、数据驱动的验证与选择循环,确保统计有效性与可部署性。在8项临床预测任务中,AgentScore超越现有评分生成方法,且在更强结构约束下达到与更灵活可解释模型相当的AUROC。在两项外部验证任务中,其判别能力高于已有指南评分。
原文摘要 · Abstract (English)
Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules. While machine-learning models achieve strong performance, many fail to translate into routine clinical use due to misalignment with workflow constraints such as memorability, auditability, and bedside execution. We argue that this gap arises not from insufficient predictive power, but from optimizing over model classes that are incompatible with guideline deployment. Deployable guidelines often take the form of unit-weighted clinical checklists, formed by thresholding the sum of binary rules, but learning such scores requires searching an exponentially large discrete space of possible rule sets. We introduce AgentScore, which performs semantically guided optimization in this space by using LLMs to propose candidate rules and a deterministic, data-grounded verification-and-selection loop to enforce statistical validity and deployability constraints. Across eight clinical prediction tasks, AgentScore outperforms existing score-generation methods and achieves AUROC comparable to more flexible interpretable models despite operating under stronger structural constraints. On two additional externally validated tasks, AgentScore achieves higher discrimination than established guideline-based scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。