arXiv:2608.20549cs.AI2026-08

让医学影像AI读懂三维数据,提升临床决策可信度

Volumetric Radiology AI in the Era of Multimodal Large Language Models

论文配图:Volumetric Radiology AI in the Era of Multimodal Large Language Models
图 1 · 摘自论文原文
  • 构建保留三维空间信息的模型表示,支持全体积分析
  • 提出可验证的智能系统框架,实现跨流程协同推理
  • 适合需要精准空间与定量分析的放射科临床场景

多模态大语言模型(MLLMs)正推动放射AI从单一图像分析迈向多模态理解与推理。然而,体积化影像存在根本性表征失配:临床解读需完整体积空间上下文及依赖采集的定量信息,而现有MLLMs多基于选定二维图像、压缩视觉表示或报告文本进行条件建模。因此,可靠的体积化放射AI需保留任务相关的三维信息,并具备跨临床流程访问、验证与整合能力。本文综述截至2026年7月超过200篇文献,围绕模型层面的体积表征与多模态理解、系统层面的代理协调机制,及其与临床应用和评估的关联展开。涵盖体积基础模型、语言对齐与压缩策略,以及通过规划、工具、记忆与工作流交互扩展的代理系统。区分了仅需二维视图或报告引导推理的场景与必须采用原生体积建模的场景。提出Claim-Design-Validation框架,评估技术、流程与临床主张是否匹配设计与验证。研究表明,原生体积建模与代理能力取决于任务的空间、定量、上下文与流程需求。临床可信度要求忠实的体积表征、可追溯的系统行为、主张一致的验证,以及在真实工作流中明确的人工监督。

原文摘要 · Abstract (English)

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.

医学影像三维建模多模态智能系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。