arXiv:2605.01779cs.CV2026-05

用迭代推理生成更准的CT报告,避免AI胡说八道。

MedScribe: Clinically Grounded CT Reporting through Agentic Workflows

论文配图:MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
图 1 · 摘自论文原文
  • 把报告生成当逐步取证过程,动态调用专用工具找病灶
  • 在两个数据集上比现有模型更准、更可信、结果更可解释
  • 适合想提升医疗影像报告质量的研究者和临床开发者

视觉语言模型(VLMs)在自动化放射科报告生成方面展现出潜力,但现有方法依赖对体数据的全局嵌入压缩,常导致虚构发现且3D CT图像解剖定位不准。我们提出MedScribe,一种基于假设驱动的框架,将报告生成重构为迭代证据获取过程,而非单次编码任务。MedScribe将报告生成建模为序列决策过程,其中大语言模型动态调用特定病理诊断工具以提取局部体数据特征。这些结构化特征用于查询与病理特定文本证据对齐的多维检索空间。通过在合成前显式累积量化证据,该框架强化了细粒度定位,并减少无依据陈述。无需任务特定微调,MedScribe在CT-RATE和RadChestCT数据集上优于最先进2D与3D VLMs,展现了假设驱动推理在可靠医学影像报告中的价值。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compression of volumetric data, often leading to hallucinated findings and limited anatomical grounding in 3D CT imaging. We introduce MedScribe, a hypothesis-driven framework that reformulates report generation as an iterative evidence acquisition process rather than a single-pass encoding task. MedScribe models reporting as a sequential decision process in which a large language model dynamically invokes pathology-specific diagnostic tools to extract localized volumetric features. These structured features are used to query a multidimensional retrieval space aligned with pathology-specific textual evidence. By explicitly accumulating quantitative evidence prior to synthesis, the framework enforces fine-grained grounding and reduces unsupported claims. Without task-specific fine-tuning, MedScribe improves clinical accuracy, factual consistency, and interpretability on CT-RATE and RadChestCT compared to state-of-the-art 2D and 3D VLMs, demonstrating the value of hypothesis-driven reasoning for reliable medical image reporting.

医学影像报告生成AI推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。