arXiv:2410.08397eess.IVcs.AI2024-10被引 7

用自然语言指令自动完成医学影像分析,省去手动调工具的麻烦。

VoxelPrompt: A Vision Agent for End-to-End Medical Image Analysis

  • 通过语言模型生成可执行代码,联动视觉网络完成分析任务。
  • 能精准测量肿瘤生长等复杂指标,准确率接近专业单任务模型。
  • 适合需要组合多种分析流程的临床研究者或医学工程师使用。

我们提出 VoxelPrompt,一个端到端的医学图像分析智能体,可处理自由格式的放射学任务。给定任意数量的三维医学影像和自然语言提示,VoxelPrompt 通过语言模型生成可执行代码,调用联合训练、可适配的视觉网络,进一步执行分析步骤以达成实际量化目标,例如跨访视测量肿瘤增长。该系统自动化了当前需人工拼接多个专用视觉与统计工具才能完成的分析流程。我们在多样化的神经影像任务中评估 VoxelPrompt,结果表明其能勾画数百个解剖与病灶特征,测量复杂的形态属性,并实现对病灶特征的开放语言分析。VoxelPrompt 在准确性上与专业单任务图像分析模型相当,同时支持广泛的组合式生物医学工作流。

原文摘要 · Abstract (English)

We present VoxelPrompt, an end-to-end image analysis agent that tackles free-form radiological tasks. Given any number of volumetric medical images and a natural language prompt, VoxelPrompt integrates a language model that generates executable code to invoke a jointly-trained, adaptable vision network. This code further carries out analytical steps to address practical quantitative aims, such as measuring the growth of a tumor across visits. The pipelines generated by VoxelPrompt automate analyses that currently require practitioners to painstakingly combine multiple specialized vision and statistical tools. We evaluate VoxelPrompt using diverse neuroimaging tasks and show that it can delineate hundreds of anatomical and pathological features, measure complex morphological properties, and perform open-language analysis of lesion characteristics. VoxelPrompt performs these objectives with an accuracy similar to that of specialist single-task models for image analysis, while facilitating a broad range of compositional biomedical workflows.

医学影像视觉代理自然语言自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。