arXiv:2505.17265cs.CLcs.AI2025-05被引 2

构建临床病例报告信息提取数据集,推动罕见病诊断AI发展

CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports

  • 构建专家标注的罕见病病例报告数据集,支持细粒度信息抽取
  • Qwen2.5-7B模型在任务中表现优于GPT-4o,类别特异性提示提升准确率
  • 验证LLM可提取临床关键信息,适合医疗AI研究者与临床辅助系统开发者

罕见病(包括先天性代谢异常,IEM)诊断困难。病例报告是重要但计算利用不足的资源。临床密集信息抽取指将医学信息结构化归类。大型语言模型(LLMs)可能实现病例报告的可扩展信息抽取,但该任务极少被评估。我们提出CaseReportBench,一个聚焦IEM的专家标注数据集,用于病例报告的密集信息抽取。基于该数据集,我们评估多种模型与提示策略,引入类别特异性提示和子标题过滤的数据整合方法。零样本链式思维提示相比标准零样本提示无显著优势。类别特异性提示提升与基准的一致性。开源模型Qwen2.5-7B在此任务上优于GPT-4o。临床医生评估显示,LLM能从病例报告中提取临床相关细节,支持罕见病诊断与管理。我们亦指出改进方向,如识别对鉴别诊断重要的阴性发现方面仍有局限。本工作推进了基于LLM的临床自然语言处理,为可扩展医疗AI应用铺路。

原文摘要 · Abstract (English)

Rare diseases, including Inborn Errors of Metabolism (IEM), pose significant diagnostic challenges. Case reports serve as key but computationally underutilized resources to inform diagnosis. Clinical dense information extraction refers to organizing medical information into structured predefined categories. Large Language Models (LLMs) may enable scalable information extraction from case reports but are rarely evaluated for this task. We introduce CaseReportBench, an expert-annotated dataset for dense information extraction of case reports, focusing on IEMs. Using this dataset, we assess various models and prompting strategies, introducing novel approaches such as category-specific prompting and subheading-filtered data integration. Zero-shot chain-of-thought prompting offers little advantage over standard zero-shot prompting. Category-specific prompting improves alignment with the benchmark. The open-source model Qwen2.5-7B outperforms GPT-4o for this task. Our clinician evaluations show that LLMs can extract clinically relevant details from case reports, supporting rare disease diagnosis and management. We also highlight areas for improvement, such as LLMs' limitations in recognizing negative findings important for differential diagnosis. This work advances LLM-driven clinical natural language processing and paves the way for scalable medical AI applications.

临床AI信息抽取罕见病LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。