让AI能像医生一样对话式解读胸片,支持多轮交互诊断。
RadVLM: A Multitask Conversational Vision-Language Model for Radiology
- 构建超百万样本的指令数据集,支持单轮与多轮对话任务。
- 在对话能力与视觉定位上达到当前最优,小样本下仍表现稳定。
- 适合临床辅助诊断场景,助力放射科医生高效工作。
胸片应用广泛但放射科医生短缺,推动了自动化分析与AI辅助报告的发展。现有视觉语言模型虽在报告生成或异常检测等任务中表现良好,但缺乏交互式诊断能力。本文提出RadVLM,一个轻量级、多任务、可对话的胸片基础模型。我们构建了一个包含超100万张图像-指令对的大规模指令数据集,涵盖单轮任务(如报告生成、异常分类、视觉定位)和多轮多任务对话。在该数据集上微调后,RadVLM在对话能力与视觉定位上达到当前最优,同时在其他放射学任务中保持竞争力。消融实验表明,多任务联合训练显著提升性能,尤其在标注数据有限时。结果表明,RadVLM具备成为临床实用AI助手的潜力,可提供结构化解读与对话支持,提升诊断效率与可及性。
原文摘要 · Abstract (English)
The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks such as report generation or abnormality detection, they often lack support for interactive diagnostic capabilities. In this work we present RadVLM, a compact, multitask conversational foundation model designed for CXR interpretation. To this end, we curate a large-scale instruction dataset comprising over 1 million image-instruction pairs containing both single-turn tasks -- such as report generation, abnormality classification, and visual grounding -- and multi-turn, multi-task conversational interactions. After fine-tuning RadVLM on this instruction dataset, we evaluate it across different tasks along with re-implemented baseline VLMs. Our results show that RadVLM achieves state-of-the-art performance in conversational capabilities and visual grounding while remaining competitive in other radiology tasks. Ablation studies further highlight the benefit of joint training across multiple tasks, particularly for scenarios with limited annotated data. Together, these findings highlight the potential of RadVLM as a clinically relevant AI assistant, providing structured CXR interpretation and conversational capabilities to support more effective and accessible diagnostic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。