用双阶段模型自动分析肺结节特征,辅助临床决策。
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

- 分两阶段:先提取影像特征,再生成描述与建议
- 结节定位准确率77.18%,直径误差仅2.58毫米
- 生成内容临床相关性高,适合放射科医生辅助使用
肺癌是全球主要癌症致死原因,计算机断层扫描(CT)是筛查和随访的主要影像工具。结节检测后,放射科医生需手动评估解剖位置、直径、边缘特征和密度类型以支持风险评估和临床决策,该流程耗时且存在观察者间差异。现有人工智能方法多聚焦单一任务,难以形成统一的临床解释框架。本研究提出FZ-VLM,一种基于微调Florence-2与Zephyr-7B的两阶段视觉语言模型框架,实现肺结节的结构化表征与临床决策支持。第一阶段使用微调后的Florence-2模型从专家标注的2D轴向CT切片中提取放射学特征,第二阶段由Zephyr-7B模型基于这些特征生成结节描述、随访建议及纵向分析。结果表明,第一阶段在解剖位置识别上达到77.18%准确率,边缘特征为67.96%,密度类型为79.13%,直径估计平均绝对误差为2.58毫米,优于对比的GPT-4基线与人类基准。第二阶段经专家评估,准确率为93.9%,完整性得分为98.6%,临床相关性为76.1%,总体评分为89.5%。安全分析显示多数输出临床安全,但部分随访建议仍需专家复核。据我们所知,这是首个用于结构化结节表征与临床决策的两阶段视觉语言模型框架。
原文摘要 · Abstract (English)
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。