arXiv:2506.09634cs.CVcs.AI2025-06被引 1

HSENet提升3D医学影像理解,解决传统方法误判问题

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

  • 双3D视觉编码器捕捉全局与局部解剖结构
  • 空间打包器压缩高分辨率3D区域,提升效率
  • 适合需要精准3D诊断的临床研究与AI辅助系统

自动化3D CT诊断通过提升诊断准确性和流程效率,助力临床做出及时、基于证据的决策。尽管多模态大语言模型在视觉-语言理解中表现优异,但现有方法主要聚焦于2D医学图像,难以捕捉复杂的3D解剖结构,常导致细微病灶误判和诊断幻觉。本文提出混合空间编码网络(HSENet),通过有效的视觉感知与投影,利用丰富的3D医学视觉线索实现精准可靠的视觉-语言理解。HSENet采用双3D视觉编码器,分别感知全局体数据上下文与精细解剖细节,并通过两阶段对齐预训练诊断报告。进一步提出空间打包器(Spatial Packer),通过基于质心的压缩,将高分辨率3D空间区域浓缩为一组信息丰富的视觉标记。结合双3D视觉编码器,HSENet可无缝将混合视觉表征传递至大语言模型语义空间,促进准确的诊断文本生成。实验表明,该方法在3D视觉-语言检索(R@100达39.85%,+5.96%)、3D医学报告生成(BLEU-4达24.01%,+8.01%)和3D视觉问答(主类准确率73.60%,+1.99%)上均达到领先性能,验证其有效性。代码已开源:https://github.com/YanzhaoShi/HSENet。

原文摘要 · Abstract (English)

Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. While multimodal large language models (MLLMs) exhibit promising performance in visual-language understanding, existing methods mainly focus on 2D medical images, which fundamentally limits their ability to capture complex 3D anatomical structures. This limitation often leads to misinterpretation of subtle pathologies and causes diagnostic hallucinations. In this paper, we present Hybrid Spatial Encoding Network (HSENet), a framework that exploits enriched 3D medical visual cues by effective visual perception and projection for accurate and robust vision-language understanding. Specifically, HSENet employs dual-3D vision encoders to perceive both global volumetric contexts and fine-grained anatomical details, which are pre-trained by dual-stage alignment with diagnostic reports. Furthermore, we propose Spatial Packer, an efficient multimodal projector that condenses high-resolution 3D spatial regions into a compact set of informative visual tokens via centroid-based compression. By assigning spatial packers with dual-3D vision encoders, HSENet can seamlessly perceive and transfer hybrid visual representations to LLM's semantic space, facilitating accurate diagnostic text generation. Experimental results demonstrate that our method achieves state-of-the-art performance in 3D language-visual retrieval (39.85% of R@100, +5.96% gain), 3D medical report generation (24.01% of BLEU-4, +8.01% gain), and 3D visual question answering (73.60% of Major Class Accuracy, +1.99% gain), confirming its effectiveness. Our code is available at https://github.com/YanzhaoShi/HSENet.

3D医学视觉语言大模型影像诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。