arXiv:2511.06752cs.CV2025-11

提出首个腹部CT症状-器官推理框架,解决多器官关联与三维结构融合难题

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

  • 用可学习的器官锚点实现症状到多个器官的软标签映射
  • 融合2D局部特征与3D全局上下文,提升推理准确率
  • 适合临床辅助诊断与医学多模态模型研究者

理解症状与影像之间的关联对临床推理至关重要。然而,现有医疗多模态模型通常依赖简单的单对一硬标签,忽略了症状可能关联多个器官的复杂现实;同时,它们主要使用单切片2D特征,缺乏3D信息,难以捕捉完整的解剖上下文。本文提出Med-SORA,一种针对腹部CT图像的症状-器官推理框架。Med-SORA引入基于RAG的数据集构建方式,采用可学习的器官锚点进行软标签标注,以建模症状到多器官的复杂关系,并设计2D-3D交叉注意力架构,融合局部与全局图像特征。据我们所知,这是首个在医疗多模态学习中系统解决症状-器官推理问题的工作。实验结果表明,Med-SORA优于现有模型,实现了更精准的三维临床推理。

原文摘要 · Abstract (English)

Understanding symptom-image associations is crucial for clinical reasoning. However, existing medical multimodal models often rely on simple one-to-one hard labeling, oversimplifying clinical reality where symptoms relate to multiple organs. In addition, they mainly use single-slice 2D features without incorporating 3D information, limiting their ability to capture full anatomical context. In this study, we propose Med-SORA, a framework for symptom-to-organ reasoning in abdominal CT images. Med-SORA introduces RAG-based dataset construction, soft labeling with learnable organ anchors to capture one-to-many symptom-organ relationships, and a 2D-3D cross-attention architecture to fuse local and global image features. To our knowledge, this is the first work to address symptom-to-organ reasoning in medical multimodal learning. Experimental results show that Med-SORA outperforms existing medical multimodal models and enables accurate 3D clinical reasoning.

医学多模态症状推理三维影像软标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。