arXiv:2609.03527cs.AI2026-09

首个专为新生儿呼吸疾病设计的多模态大模型,提升诊断报告准确性。

NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

论文配图:NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis
图 1 · 摘自论文原文
  • 构建知识-逻辑-对齐框架,融合新生儿临床先验与影像逻辑
  • 在新生儿数据集上达53.29% ROUGE-L与65.19%临床效能F1
  • 既适配新生儿诊断,也保持成人影像报告生成能力

新生儿呼吸系统疾病是导致新生儿发病率和死亡率的主要原因,临床实践中面临巨大挑战。现有多模态大模型(MLLMs)在新生儿诊断中存在两大局限:(1)主要基于成人数据训练导致领域差距;(2)缺乏多维临床信息的整合。为此,我们收集了两个真实临床数据集(NeoCXR 和 NeoCXR-EV),提出 NeoRed,据我们所知,首个专为新生儿呼吸疾病诊断设计的 MLLM,填补了新生儿诊断报告生成的空白。为增强异构临床信息与胸部 X 光片的联合诊断能力,我们设计了新颖的知识-逻辑-对齐(KLA)框架,从三方面约束模型行为:1)知识先验注入(KPI)将新生儿科医生启发的诊断先验嵌入多模态表示,引导跨模态疾病特异性注意力;2)诊断逻辑约束(DLC)使生成报告语义与多模态诊断逻辑对齐;3)视觉语义对齐(VSA)建立视觉特征与影像结论间的语义对应。大量实验表明,NeoRed 能够实现精准的新生儿诊断报告生成,在 NeoCXR 上达到 53.29% 的 ROUGE-L 和 65.19% 的临床效能 F1,优于现有 MLLMs。NeoRed 在成人基准数据集(MIMIC-CXR 与 IU-Xray)上也保持了竞争力。数据集申请后开放。

原文摘要 · Abstract (English)

Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain gap arising from predominantly adult training data; (2) insufficient integration of multidimensional clinical context for accurate diagnosis. To address these challenges, we collect two real-world clinical datasets (NeoCXR and NeoCXR-EV) and propose NeoRed, to the best of our knowledge, the first MLLM tailored for neonatal respiratory disease, filling the gap in neonatal diagnostic reports generation. To enhance joint diagnosis from heterogeneous clinical context and chest X-rays, we design a novel Knowledge-Logic-Alignment (KLA) framework which constrains model behavior from three perspectives: 1) Knowledge Prior Injection (KPI) incorporates neonatologist-inspired diagnostic priors into multimodal representations, guiding disease-specific attention across modalities; 2) Diagnostic Logic Constraint (DLC) aligns the semantics of generated reports with multimodal diagnostic logic; and 3) Visual Semantic Alignment (VSA) establishes semantic correspondence between visual features and imaging conclusions. Extensive experiments demonstrate that NeoRed enables accurate neonatal diagnostic reports generation, achieving ROUGE-L of 53.29% and Clinical Efficacy F1 score of 65.19% on NeoCXR, outperforming existing MLLMs. NeoRed also preserves competitive report generation performance on adult benchmarks (MIMIC-CXR and IU-Xray). Datasets will be available upon application.

新生儿诊断多模态大模型医学AIX光分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。