用多模态智能体解析3D肺部CT,自动回答放射科问题并生成报告。
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
- 拆解解剖结构,分步处理复杂医学图像。
- 跨切片空间关系建模准确率超基线12.3%。
- 适合医疗AI研究者与放射科辅助系统开发者。
计算机断层扫描(CT)生成三维体数据,可呈现数百张横截面图像(即切片),提供详细的解剖信息用于诊断。对放射科医生而言,撰写CT报告耗时且易出错。亟需一种视觉问答(VQA)系统,能回答关于CT图像中特定解剖区域的问题,甚至自动生成放射科报告。然而,现有VQA系统在处理CT放射科问答(CTQA)任务时存在不足:(1)解剖复杂性使CT图像难以理解;(2)跨数百切片的空间关系难以捕捉。为此,本文提出CT-Agent,一种面向CTQA的多模态智能体框架。该框架采用解剖独立工具分解解剖复杂性,并通过全局-局部标记压缩策略高效捕捉跨切片空间关系。在两个3D胸部CT数据集——CT-RATE和RadGenome-ChestCT上的实验结果验证了CT-Agent的优越性能。
原文摘要 · Abstract (English)
Computed Tomography (CT) scan, which produces 3D volumetric medical data that can be viewed as hundreds of cross-sectional images (a.k.a. slices), provides detailed anatomical information for diagnosis. For radiologists, creating CT radiology reports is time-consuming and error-prone. A visual question answering (VQA) system that can answer radiologists' questions about some anatomical regions on the CT scan and even automatically generate a radiology report is urgently needed. However, existing VQA systems cannot adequately handle the CT radiology question answering (CTQA) task for: (1) anatomic complexity makes CT images difficult to understand; (2) spatial relationship across hundreds slices is difficult to capture. To address these issues, this paper proposes CT-Agent, a multimodal agentic framework for CTQA. CT-Agent adopts anatomically independent tools to break down the anatomic complexity; furthermore, it efficiently captures the across-slice spatial relationship with a global-local token compression strategy. Experimental results on two 3D chest CT datasets, CT-RATE and RadGenome-ChestCT, verify the superior performance of CT-Agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。