构建首个低资源语言3D PET/CT细粒度病灶标注数据集并提出临床流程对齐生成框架。
Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

- 基于图结构建模病灶间关系,模拟放射科医生局部分析诊断流程。
- 在新指标上提升45.8%临床可靠性,生成报告更贴近真实病灶描述。
- 适用于医学影像智能报告生成、低资源语言医疗AI研究者。
3D PET/CT自动报告生成面临高维体数据与标注数据稀缺的双重挑战,尤其在低资源语言场景下。现有黑箱方法将整体积映射为报告,忽略临床中逐区域病灶分析的诊断流程。本文提出VietPET-RoI,首个面向低资源语言的大规模3D PET/CT细粒度病灶标注数据集,包含600例扫描样本与1,960个手动标注的病灶区域(RoIs),对应临床报告。为进一步利用该数据集,提出HiRRA框架,通过图神经网络捕捉病灶属性间的关联,实现从全局模式匹配到局部临床发现的转变。引入新评估指标RoI Coverage与RoI Quality Index,基于LLM提取衡量病灶定位准确率与特征描述忠实度。大量实验表明,该框架在BLEU上超越现有模型19.7%,在ROUGE-L上提升4.7%,临床指标提升45.8%,显著增强报告可靠性并减少幻觉。代码与数据集已开源。
原文摘要 · Abstract (English)
Automated medical report generation for 3D PET/CT imaging is fundamentally challenged by the high-dimensional nature of volumetric data and a critical scarcity of annotated datasets, particularly for low-resource languages. Current black-box methods map whole volumes to reports, ignoring the clinical workflow of analyzing localized Regions of Interest (RoIs) to derive diagnostic conclusions. In this paper, we bridge this gap by introducing VietPET-RoI, the first large-scale 3D PET/CT dataset with fine-grained RoI annotation for a low-resource language, comprising 600 PET/CT samples and 1,960 manually annotated RoIs, paired with corresponding clinical reports. Furthermore, to demonstrate the utility of this dataset, we propose HiRRA, a novel framework that mimics the professional radiologist diagnostic workflow by employing graph-based relational modules to capture dependencies between RoI attributes. This approach shifts from global pattern matching toward localized clinical findings. Additionally, we introduce new clinical evaluation metrics, namely RoI Coverage and RoI Quality Index, that measure both RoI localization accuracy and attribute description fidelity using LLM-based extraction. Extensive evaluation demonstrates that our framework achieves SOTA performance, surpassing existing models by 19.7% in BLEU and 4.7% in ROUGE-L, while achieving a remarkable 45.8% improvement in clinical metrics, indicating enhanced clinical reliability and reduced hallucination. Our code and dataset are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。