构建呼吸科临床决策基准,评估大模型在慢阻肺与肺结节管理中的表现。
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

- 基于真实临床数据构建多模态决策评估基准
- 大模型平均得分68.58,最高达72.48,存在影像幻觉与用药风险
- 适合医疗AI研发者、临床医生及模型安全评估人员使用
呼吸专科诊疗需多模态分析、长期风险评估、遵循指南的干预及全程管理,现有医学评测难以覆盖。本文构建了基于真实场景的RESPClinBench基准,用于评估大语言模型在呼吸科临床决策中的表现。基准包含427例开放性慢阻肺急性加重(AECOPD-PIM)病例和196例结合胸部CT与结构化信息的肺结节(PNBIM)病例。七种大模型通过标准化API推理生成4,361条响应,采用原子动作召回率与人工评分框架综合打分。结果显示,总平均得分为68.58,其中Qwen3.6-27B整体排名第一(71.22),在PNBIM中以72.48领先,于AECOPD-PIM中同样以71.11位居第一。在PNBIM中,31.85%的回复出现影像幻觉,8.16%存在严重医疗风险;在AECOPD-PIM中,26.93%存在用药安全风险,1.44%存在严重医疗风险。研究证明,该基准可揭示模型在多模态肺结节评估与慢阻肺管理中的任务特异性局限,并通过临床动作覆盖率、整体评估与独立安全预警机制,为模型选型与前瞻性验证提供可靠依据。
原文摘要 · Abstract (English)
Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。