arXiv:2605.25357cs.CVcs.MA2026-05

用多智能体协作提升胎儿超声分析的可靠性与可解释性

Towards Reliable Fetal Ultrasound Interpretation with Multi-Agent Collaboration

论文配图:Towards Reliable Fetal Ultrasound Interpretation with Multi-Agent Collaboration
图 1 · 摘自论文原文
  • 通过多智能体分工协作,分解临床问题为识别-测量-报告流程
  • 在胎儿超声VQA任务上比最强基线高出25%以上准确率
  • 适合产科影像医生、医学AI研发者参考,推动临床辅助系统落地

自动化胎儿超声解读需完成从图像感知(平面识别、解剖分割)到临床理解(生物测量、诊断报告)的全流程。现有“一任务一模型”范式难以整合多步骤证据。尽管多模态大语言模型(MLLMs)具备良好视觉理解能力,但其领域特定知识不足且易产生幻觉,限制了在胎儿超声分析中的可靠性。为此,我们提出FetUSAgents——一种工具增强的多智能体系统,支持视觉问答(VQA)、报告生成、图像描述和视频摘要。该系统通过协作式LLM智能体协调专用视觉工具,将临床查询分解为从解剖识别到定量测量的子任务。我们进一步引入双路径证据仲裁(DPEA),融合基于LLM的推理与专用视觉工具的结构化证据。检索增强的证据库整合中间结果,确保结论可追溯、临床可信赖。此外,我们构建了FetUS-VQA基准,包含1,892张图像和3,205个跨10项临床任务的问答对。大量分布外实验表明,FetUSAgents在VQA准确率上超越通用与医学MLLMs,优于最强基线超过25%。结果表明,这是一条可扩展的证据驱动型产前影像临床助手路径。代码已开源。

原文摘要 · Abstract (English)

Automated fetal ultrasound interpretation requires a workflow from visual perception, including plane recognition and anatomical segmentation, to clinical understanding, including biometric measurement and diagnostic reporting. However, the prevailing "one-task, one-model" paradigm limits systematic integration of evidence across this multi-step process. Although multimodal large language models (MLLMs) show promising visual understanding, their limited domain-specific grounding and hallucination risks restrict reliability in fetal ultrasound analysis. To address these limitations, we propose FetUSAgents, a tool-augmented multi-agent system for comprehensive fetal ultrasound interpretation, supporting visual question answering (VQA), report generation, image captioning, and video summarization. FetUSAgents coordinates task-specific visual tools through collaborative LLM agents and decomposes clinical queries into subtasks that progress from anatomical recognition to quantitative measurement. We further introduce Dual-Path Evidence Arbitration (DPEA), which integrates LLM-based deliberative reasoning with structured computational evidence from specialized visual tools. A retrieval-enhanced evidence bank consolidates intermediate findings to support traceable and clinically grounded conclusions. In addition, we construct FetUS-VQA, a dedicated VQA benchmark for fetal ultrasound, comprising 1,892 images and 3,205 question-answer pairs across 10 clinical tasks. Extensive out-of-distribution experiments show that FetUSAgents outperforms general and medical MLLMs, exceeding the strongest baseline by more than 25 percent in VQA accuracy. These results suggest a scalable route toward evidence-driven clinical assistants for prenatal imaging. Code is available.

医学影像多智能体超声分析LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。