构建临床文本转SQL新基准,推动真实电子病历智能分析
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
- 设计多表关联与患者相似队列推理的复杂查询任务
- 主流模型在困难任务上执行率仅67%-75%,距临床可用仍有差距
- 适合医疗AI、自然语言接口研究者关注临床可靠性提升
真实世界临床文本转SQL需处理异构电子病历表格、时间窗口及患者相似队列,以生成可执行查询。我们提出CLINSQL基准,包含633个专家标注任务,基于MIMIC-IV v3.1,要求多表连接、临床有意义过滤和可执行SQL。解决该任务需理解模式元数据与临床编码系统,处理长上下文,并生成超越传统文本转SQL的多步查询。我们评估22个专有与开源模型,采用链式思维自精炼策略,并通过带执行检查的评分体系,优先保障关键临床需求。尽管近期进展显著,性能仍远未达临床可靠性:测试集上,GPT-5-mini得分为74.7%,DeepSeek-R1为69.2%(开源最优),Gemini-2.5-Pro在难例上从85.5%降至67.2%。CLINSQL的进步标志着向真实电子病历分析中临床可靠文本转SQL的重要迈进。
原文摘要 · Abstract (English)
Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving CLINSQL entails navigating schema metadata and clinical coding systems, handling long contexts, and composing multi-step queries beyond traditional text-to-SQL. We evaluate 22 proprietary and open-source models under Chain-of-Thought self-refinement and use rubric-based SQL analysis with execution checks that prioritize critical clinical requirements. Despite recent advances, performance remains far from clinical reliability: on the test set, GPT-5-mini attains 74.7% execution score, DeepSeek-R1 leads open-source at 69.2% and Gemini-2.5-Pro drops from 85.5% on Easy to 67.2% on Hard. Progress on CLINSQL marks tangible advances toward clinically reliable text-to-SQL for real-world EHR analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。