让胸部X光诊断更可靠:用证据驱动的AI助手避免胡说八道
CXReasonAgent: Evidence-Grounded Diagnostic Reasoning Agent for Chest X-rays
- 用大语言模型+临床工具结合,让诊断推理有图像证据支持
- 在12个任务上生成1946轮对话,回答比大视觉语言模型更可信
- 适合医疗AI研发者和需要可验证诊断的临床场景
胸部X光在胸腔疾病诊断中至关重要,其解读需要多步骤、基于证据的推理。然而,大型视觉语言模型(LVLM)常生成看似合理但缺乏真实影像证据支撑的回答,且难以验证,还需昂贵重训练以适应新任务,限制了其在临床中的可靠性与适应性。为此,我们提出CXReasonAgent,一个将大语言模型(LLM)与临床诊断工具结合的诊断智能体,利用图像提取的诊断与视觉证据进行证据锚定的推理。为评估该能力,我们构建了包含1,946轮对话的多轮对话基准数据集CXReasonDial,涵盖12项诊断任务。实验表明,与LVLM相比,CXReasonAgent能生成更忠实于证据的回答,实现更可靠、可验证的诊断推理。结果强调了在安全关键的临床场景中整合临床锚定工具的重要性。演示链接:https://ttumyche.github.io/cxreasonagent/#demo
原文摘要 · Abstract (English)
Chest X-ray plays a central role in thoracic diagnosis, and its interpretation inherently requires multi-step, evidence-grounded reasoning. However, large vision-language models (LVLMs) often generate plausible responses that are not faithfully grounded in diagnostic evidence and provide limited visual evidence for verification, while also requiring costly retraining to support new diagnostic tasks, limiting their reliability and adaptability in clinical settings. To address these limitations, we present CXReasonAgent, a diagnostic agent that integrates a large language model (LLM) with clinically grounded diagnostic tools to perform evidence-grounded diagnostic reasoning using image-derived diagnostic and visual evidence. To evaluate these capabilities, we introduce CXReasonDial, a multi-turn dialogue benchmark with 1,946 dialogues across 12 diagnostic tasks, and show that CXReasonAgent produces faithfully grounded responses, enabling more reliable and verifiable diagnostic reasoning than LVLMs. These findings highlight the importance of integrating clinically grounded diagnostic tools, particularly in safety-critical clinical settings. The demo is available \href{https://ttumyche.github.io/cxreasonagent/#demo}{here}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。