用概念瓶颈+多智能体RAG提升X光报告可解释性
Towards Interpretable Radiology Report Generation via Concept Bottlenecks using a Multi-Agentic RAG
- 通过概念瓶颈建模视觉特征与临床概念关系
- 在COVID-QU数据集上达81%分类准确率,报告生成五项指标84%-90%
- 适合需要可解释AI辅助诊断的临床场景
深度学习虽推动医学图像分类进步,但可解释性不足限制其临床应用。本研究通过概念瓶颈模型(CBMs)与多智能体检索增强生成(RAG)系统,提升胸部X光(CXR)分类的可解释性。通过建模视觉特征与临床概念间的关联,生成可解释的概念向量,指导多智能体RAG系统生成具有临床相关性、可解释性和透明性的放射科报告。利用大模型作为评判者评估生成报告,验证了模型输出的可解释性与临床价值。在COVID-QU数据集上,模型分类准确率达81%,报告生成性能稳健,五项关键指标介于84%至90%之间。该可解释的多智能体框架弥合了高性能AI与临床所需可解释性之间的差距,适用于可靠的AI驱动胸部X光分析。代码已开源:https://github.com/tifat58/IRR-with-CBM-RAG.git。
原文摘要 · Abstract (English)
Deep learning has advanced medical image classification, but interpretability challenges hinder its clinical adoption. This study enhances interpretability in Chest X-ray (CXR) classification by using concept bottleneck models (CBMs) and a multi-agent Retrieval-Augmented Generation (RAG) system for report generation. By modeling relationships between visual features and clinical concepts, we create interpretable concept vectors that guide a multi-agent RAG system to generate radiology reports, enhancing clinical relevance, explainability, and transparency. Evaluation of the generated reports using an LLM-as-a-judge confirmed the interpretability and clinical utility of our model's outputs. On the COVID-QU dataset, our model achieved 81% classification accuracy and demonstrated robust report generation performance, with five key metrics ranging between 84% and 90%. This interpretable multi-agent framework bridges the gap between high-performance AI and the explainability required for reliable AI-driven CXR analysis in clinical settings. Our code is available at https://github.com/tifat58/IRR-with-CBM-RAG.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。