arXiv:2509.24739cs.CV2025-09NeurIPS被引 11

首个越南语PET/CT报告生成数据集,助力低资源语言医疗AI发展

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

  • 构建2757例越南语PET/CT影像与报告配对数据集
  • 在下游任务中显著提升现有视觉语言模型性能
  • 填补低资源语言医疗影像数据空白,适合临床应用研究

视觉-语言基础模型(VLMs)在大规模多模态数据上训练后,在人工智能领域推动了跨模态推理的进展。然而,由于医学影像数据多样性和多语言临床数据稀缺,将其应用于医学影像仍面临挑战。现有医学VLMs多局限于部分成像模态且集中于高资源语言,限制了其泛化能力和临床实用性。为此,我们引入一个全新的越南语多模态医学数据集,包含2,757例独立患者的全身PET/CT影像及其完整临床报告。该数据集旨在填补两大空白:(1)现有VLM训练语料中缺乏PET/CT影像数据,阻碍功能成像任务模型的发展;(2)低资源语言(特别是越南语)在医学视觉-语言研究中的代表性不足。据我们所知,这是首个提供全面越南语PET/CT-报告配对的数据集。我们进一步提出训练框架,包含数据增强和专家验证测试集。通过全面实验,在下游任务上评估主流VLMs表现,结果表明引入本数据集可显著提升模型性能。我们认为该数据集与基准将为推动更鲁棒的医学影像VLM发展,尤其在低资源语言和越南医疗场景中发挥关键作用。源代码见:https://github.com/AIoT-Lab-BKAI/ViPET-ReportGen。

原文摘要 · Abstract (English)

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, especially for low-resource languages and clinical use in Vietnamese healthcare. The source code is available at https://github.com/AIoT-Lab-BKAI/ViPET-ReportGen.

医疗AI多模态越南语PET/CT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。