用感知生成的事实增强提示,显著提升3D脑MRI报告质量
PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

- 将3D分割结果转为结构化事实句注入提示
- 事实注入使报告质量超越检索旧报告,且无需真实标注
- 适合医学影像生成、提示工程研究者参考
放射科报告生成已高度依赖2D胸部X光片,通常通过更大模型或医疗数据预训练提升性能。本文在3D多序列脑部MRI这一体积分割多疾病场景下重新审视该假设,发现模型大小并非关键。零样本医疗视觉语言模型在脑部MRI上表现不佳,胸部影像专用模型尤其失败;五个不同规模的骨干网络在三类模型家族中微调后,性能差异微小。真正决定报告质量的是提示中注入的信息。我们采用上游3D分割与分类结果,将其序列化为结构化事实句,通过LoRA适配的视觉语言模型进行提示,提出方法PerFact。在控制变量实验中,感知生成的事实优于检索历史报告,一旦事实存在,检索即冗余;端到端预测的事实在无真实标注时仍有效。预测事实与理想事实间的残差差距源于事实粒度,而非生成器能力。封闭式视觉问答对报告质量无负面影响,但其信息源影响甚微。在3D脑部MRI中,可控制因素中,信息接地性远超模型选择,是决定报告质量的核心。
原文摘要 · Abstract (English)
Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D multi-sequence brain MRI, a volumetric multi-disease regime, and find that the model is not the lever. Zero-shot medical and radiology vision-language models transfer poorly to brain MRI, with chest radiograph specialists failing most conspicuously, and five backbones fine-tuned identically across three model families and an order of magnitude in scale differ only marginally. What determines the quality of the report is the information injected into the prompt. We delegate perception to upstream 3D segmentation and classification, serialize their outputs into a structured fact sentence, and prompt a LoRA-adapted vision-language model with it; we call this \textbf{PerFact}. In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth annotation at inference. The residual gap between predicted and oracle facts is explained by the granularity of the facts rather than by the generator. Closed-ended visual question answering comes at no measurable cost to report quality, though the grounding source has little effect on it. On 3D brain MRI, grounding information, not model choice, is the dominant controllable factor in report quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。