arXiv:2504.11008cs.CVcs.AI2025-04被引 13

让AI理解口语化医学提问并精准定位病灶

MediSee: Reasoning-based Pixel-level Perception in Medical Images

  • 基于多视角逻辑推理构建医疗问答数据集
  • 支持口语化提问,准确生成病灶分割与框选
  • 适合临床辅助诊断与普通用户交互场景

尽管像素级医学图像感知取得显著进展,现有方法或局限于特定任务,或严重依赖精确的边界框或文本标签作为输入提示。然而,这些专业信息对普通用户构成巨大障碍,极大限制了方法的通用性。相较而言,普通用户更倾向于使用需要逻辑推理的口头查询。本文提出一项新型医学视觉任务:医学推理分割与检测(MedSD),旨在理解关于医学图像的隐含查询,并为目标对象生成对应的分割掩码和边界框。为此,我们首先构建了多视角、逻辑驱动的医学推理分割与检测(MLMR-SD)数据集,包含大量医学实体目标及其对应推理过程。此外,我们提出了MediSee模型,作为该任务的有效基线。实验表明,该方法能有效应对包含非正式口语提问的MedSD任务,优于传统医学指代分割方法。

原文摘要 · Abstract (English)

Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge required for input is a huge obstacle for general public, which greatly reduces the universality of these methods. Compared with these domain-specialized auxiliary information, general users tend to rely on oral queries that require logical reasoning. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for medical reasoning segmentation and detection. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods.

医学图像推理分割自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。