用多模态模型自动诊断肋骨骨折并生成报告,准确率超人类专家。
OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models
- 融合CT图像与医学知识图谱,用YOLOv9和LLaVA实现智能诊断。
- 在2.8万张标注CT上平均得分4.28,优于GPT-4和Claude-3。
- 适合放射科医生辅助阅片,提升诊断效率与报告质量。
医疗影像数据量激增,对骨骼损伤如肋骨骨折的自动化诊断需求迫切,常规依赖CT扫描。人工判读耗时且易出错。本文提出OrthoInsight,一种基于多模态深度学习的肋骨骨折诊断与报告生成框架。该系统集成YOLOv9用于骨折检测、医学知识图谱获取临床背景信息,并采用微调后的LLaVA语言模型生成诊断报告。通过结合CT图像的视觉特征与专家文本数据,输出具有临床价值的结果。在28,675张标注CT图像及专家报告上评估,其诊断准确性、内容完整性、逻辑连贯性与临床指导价值四项指标平均得分达4.28,优于GPT-4与Claude-3。本研究展示了多模态学习在医学影像分析中的潜力,可为放射科医生提供有效支持。
原文摘要 · Abstract (English)
The growing volume of medical imaging data has increased the need for automated diagnostic tools, especially for musculoskeletal injuries like rib fractures, commonly detected via CT scans. Manual interpretation is time-consuming and error-prone. We propose OrthoInsight, a multi-modal deep learning framework for rib fracture diagnosis and report generation. It integrates a YOLOv9 model for fracture detection, a medical knowledge graph for retrieving clinical context, and a fine-tuned LLaVA language model for generating diagnostic reports. OrthoInsight combines visual features from CT images with expert textual data to deliver clinically useful outputs. Evaluated on 28,675 annotated CT images and expert reports, it achieves high performance across Diagnostic Accuracy, Content Completeness, Logical Coherence, and Clinical Guidance Value, with an average score of 4.28, outperforming models like GPT-4 and Claude-3. This study demonstrates the potential of multi-modal learning in transforming medical image analysis and providing effective support for radiologists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。