AURA让AI能看懂医学影像并解释推理过程,实现交互式诊断支持。
AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
- 构建多模态医疗智能体,融合图像分割与反事实生成能力。
- 通过像素级差异图等工具评估诊断相关性与可解释性。
- 适合临床辅助决策、AI可解释性研究者使用。
大型语言模型(LLMs)的进展推动了从静态预测系统向具备推理、工具调用和任务适应能力的智能体范式的转变。尽管基于LLM的智能体在多个领域展现出潜力,其在医学影像中的应用仍处于初级阶段。本文提出AURA,首个专为医学影像全面分析、解释与评估设计的视觉-语言可解释性智能体。通过动态交互、上下文解释和假设验证,AURA显著提升了AI系统的透明度、适应性和临床契合度。依托Qwen-32B架构,AURA集成模块化工具箱:(i) 分割套件,包含相位定位、病灶分割与解剖结构分割,用于定位临床意义区域;(ii) 反事实图像生成模块,支持图像级解释与推理;(iii) 评估工具集,包括像素级差异图分析、分类及前沿组件,用于评估诊断相关性与视觉可解释性。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have catalyzed a paradigm shift from static prediction systems to agentic AI agents capable of reasoning, interacting with tools, and adapting to complex tasks. While LLM-based agentic systems have shown promise across many domains, their application to medical imaging remains in its infancy. In this work, we introduce AURA, the first visual linguistic explainability agent designed specifically for comprehensive analysis, explanation, and evaluation of medical images. By enabling dynamic interactions, contextual explanations, and hypothesis testing, AURA represents a significant advancement toward more transparent, adaptable, and clinically aligned AI systems. We highlight the promise of agentic AI in transforming medical image analysis from static predictions to interactive decision support. Leveraging Qwen-32B, an LLM-based architecture, AURA integrates a modular toolbox comprising: (i) a segmentation suite with phase grounding, pathology segmentation, and anatomy segmentation to localize clinically meaningful regions; (ii) a counterfactual image-generation module that supports reasoning through image-level explanations; and (iii) a set of evaluation tools including pixel-wise difference-map analysis, classification, and advanced state-of-the-art components to assess diagnostic relevance and visual interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。