arXiv:2603.15513cs.CL2026-03

首个越南胸部X光多模态数据集,助力AI更准确诊断本地患者。

ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models

  • 构建5400张越南胸片与专家报告配对数据集
  • 模型在越南医学语境下易产生幻觉,精度不高
  • 为本地医疗AI研究提供基准,适合临床AI开发者

越南医疗研究日益重要,尤其在智能技术减轻临床诊断时间与资源负担方面。尽管视觉语言模型(如Gemini和GPT-4V)取得进展,但多数模型缺乏越南医学数据训练,难以生成符合越南患者情境的准确诊断输出。为此,我们推出ViX-Ray,一个包含5,400张越南胸片的全新数据集,每张图像均配有来自越南大型医院放射科医生撰写的影像发现与诊断意见。我们分析了报告中的语言模式,包括解剖部位与诊断术语的出现频率,揭示越南放射报告的领域特异性语言特征。进一步地,我们在ViX-Ray上微调五种先进开源视觉语言模型,并与GPT-4V和Gemini等主流闭源模型对比。结果显示,尽管部分模型输出与临床真实情况部分吻合,但在生成诊断意见时普遍存在低精度与严重幻觉问题。这些发现不仅体现了该数据集的复杂性与挑战性,也确立了ViX-Ray作为评估和推动越南临床领域视觉语言模型发展的关键基准。

原文摘要 · Abstract (English)

Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs), such as Gemini and GPT-4V, have sparked a growing interest in applying AI to healthcare. However, most existing VLMs lack exposure to Vietnamese medical data, limiting their ability to generate accurate and contextually appropriate diagnostic outputs for Vietnamese patients. To address this challenge, we introduce ViX-Ray, a novel dataset comprising 5,400 Vietnamese chest X-ray images annotated with expert-written findings and impressions from physicians at a major Vietnamese hospital. We analyze linguistic patterns within the dataset, including the frequency of mentioned body parts and diagnoses, to identify domain-specific linguistic characteristics of Vietnamese radiology reports. Furthermore, we fine-tune five state-of-the-art open-source VLMs on ViX-Ray and compare their performance to leading proprietary models, GPT-4V and Gemini. Our results show that while several models generate outputs partially aligned with clinical ground truths, they often suffer from low precision and excessive hallucination, especially in impression generation. These findings not only demonstrate the complexity and challenge of our dataset but also establish ViX-Ray as a valuable benchmark for evaluating and advancing vision-language models in the Vietnamese clinical domain.

医学影像多模态越南语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。