arXiv:2507.04383eess.IVcs.CV2025-07

构建多模态卵巢肿瘤数据集与分类模型,提升早期诊断准确率

ViTaL: A Multimodality Dataset and Benchmark for Multi-pathological Ovarian Tumor Recognition

  • 整合影像、表格和报告三类数据,构建496例患者多模态数据集
  • 提出基于三重层级偏移注意力的ViTaL-Net,多病种分类准确率超90%
  • 适用于医学AI研究者,尤其关注妇科肿瘤智能诊断方向

卵巢肿瘤是常见妇科疾病,早期未发现可迅速恶化,威胁女性健康。深度神经网络有望识别卵巢肿瘤以降低死亡率,但公开数据集有限制约了发展。为此,我们推出名为 extbf{ViTaL}的多模态病理识别数据集,涵盖496名患者的视觉、表格和语言模态数据,分为六类病理类型。该数据集包含2216张二维超声图像、496例患者的体检表格数据及496份超声报告文本。临床中仅区分良恶性不足,为实现多病种分类,我们提出基于三重层级偏移注意力机制(THOAM)的ViTaL-Net,有效减少多模态特征融合中的损失,增强跨模态信息相关性与互补性。在全面实验中,该方法在两种最常见病理类型上准确率超过90%,整体性能达85%。数据与代码已开源:https://github.com/GGbond-study/vitalnet。

原文摘要 · Abstract (English)

Ovarian tumor, as a common gynecological disease, can rapidly deteriorate into serious health crises when undetected early, thus posing significant threats to the health of women. Deep neural networks have the potential to identify ovarian tumors, thereby reducing mortality rates, but limited public datasets hinder its progress. To address this gap, we introduce a vital ovarian tumor pathological recognition dataset called \textbf{ViTaL} that contains \textbf{V}isual, \textbf{T}abular and \textbf{L}inguistic modality data of 496 patients across six pathological categories. The ViTaL dataset comprises three subsets corresponding to different patient data modalities: visual data from 2216 two-dimensional ultrasound images, tabular data from medical examinations of 496 patients, and linguistic data from ultrasound reports of 496 patients. It is insufficient to merely distinguish between benign and malignant ovarian tumors in clinical practice. To enable multi-pathology classification of ovarian tumor, we propose a ViTaL-Net based on the Triplet Hierarchical Offset Attention Mechanism (THOAM) to minimize the loss incurred during feature fusion of multi-modal data. This mechanism could effectively enhance the relevance and complementarity between information from different modalities. ViTaL-Net serves as a benchmark for the task of multi-pathology, multi-modality classification of ovarian tumors. In our comprehensive experiments, the proposed method exhibited satisfactory performance, achieving accuracies exceeding 90\% on the two most common pathological types of ovarian tumor and an overall performance of 85\%. Our dataset and code are available at https://github.com/GGbond-study/vitalnet.

多模态学习医学影像卵巢肿瘤分类模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。