用异常中心思维让视频大模型精准生成肺部CT报告
Unleashing Video Language Models for Fine-grained HRCT Report Generation
- 以病灶为中心设计推理链,引导模型关注关键异常
- 通过临床易混淆病灶作为难例优化,提升细粒度识别能力
- 在不依赖专用医学预训练的情况下超越专业模型
从高分辨率计算机断层扫描(HRCT)生成精确诊断报告对临床工作流至关重要,但受三维数据中病理多样性高、空间稀疏性强的影响,仍具挑战。尽管视频语言模型(VideoLMs)在通用领域展现强大时空推理能力,其在高体积医学图像这一特定领域的适应性尚未充分探索。本文提出AbSteering框架,一种以异常为中心的范式,引导通用视频大模型生成精细的HRCT报告。该框架包含:(i) 异常中心的思维链机制,强制模型进行病灶推理;(ii) 直接偏好优化目标,利用临床易混淆异常作为硬负样本,增强细粒度区分能力。实验表明,经此引导后,通用视频大模型在高体积医学影像任务中具备强迁移能力。值得注意的是,AbSteering在检测敏感性上优于当前最先进的专用CT基础模型(这些模型基于大规模CT数据预训练),同时有效减少幻觉现象。
原文摘要 · Abstract (English)
Generating precise diagnostic reports from High-Resolution Computed Tomography (HRCT) is critical for clinical workflow, yet it remains a formidable challenge due to the high pathological diversity and spatial sparsity within 3D volumes. While Video Language Models (VideoLMs) have demonstrated remarkable spatio-temporal reasoning in general domains, their adaptability to domain-specific, high-volume medical interpretation remains underexplored. In this work, we present AbSteering, an abnormality-centric framework that steers VideoLMs toward precise HRCT report generation. Specifically, AbSteering introduces: (i) an abnormality-centric Chain-of-Thought scheme that enforces abnormality reasoning, and (ii) a Direct Preference Optimization objective that utilizes clinically confusable abnormalities as hard negatives to enhance fine-grained discrimination. Our results demonstrate that general-purpose VideoLMs possess strong transferability to high-volume medical imaging when guided by this paradigm. Notably, AbSteering outperforms state-of-the-art domain-specific CT foundation models, which are pretrained with large-scale CTs, achieving superior detection sensitivity while simultaneously mitigating hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。