首个胎儿超声多智能体系统,实现全流程自动化分析与报告生成。
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
- 用多个专业视觉智能体动态协作,统一完成诊断、测量、分割任务。
- 在八项临床任务中表现优于专用模型和多模态大模型,准确率领先。
- 支持视频流自动摘要与结构化报告生成,适合产科临床落地使用。
胎儿超声是产前筛查的主要影像手段,但其解读高度依赖临床医生经验。尽管深度学习和基础模型取得进展,现有自动化工具仍难以兼顾任务特定精度与全流程泛化能力,无法满足端到端临床工作流需求。为此,我们提出FetalAgents——首个面向胎儿超声图像与视频分析的多智能体系统。通过轻量级代理协调框架,FetalAgents动态调度多个专业化视觉专家,最大化诊断、测量与分割任务的表现。此外,该系统突破静态图像分析局限,支持端到端视频流摘要:自动识别多解剖平面的关键帧,由协同专家分析,并融合患者元数据生成结构化临床报告。在八个临床任务上的多中心外部评估显示,FetalAgents在性能鲁棒性与准确性上均优于专用模型及多模态大语言模型(MLLMs),最终提供可审计、流程对齐的胎儿超声分析与报告解决方案。
原文摘要 · Abstract (English)
Fetal ultrasound (US) is the primary imaging modality for prenatal screening, yet its interpretation relies heavily on the expertise of the clinician. Despite advances in deep learning and foundation models, existing automated tools for fetal US analysis struggle to balance task-specific accuracy with the whole-process versatility required to support end-to-end clinical workflows. To address these limitations, we propose FetalAgents, the first multi-agent system for comprehensive fetal US analysis. Through a lightweight, agentic coordination framework, FetalAgents dynamically orchestrates specialized vision experts to maximize performance across diagnosis, measurement, and segmentation. Furthermore, FetalAgents advances beyond static image analysis by supporting end-to-end video stream summarization, where keyframes are automatically identified across multiple anatomical planes, analyzed by coordinated experts, and synthesized with patient metadata into a structured clinical report. Extensive multi-center external evaluations across eight clinical tasks demonstrate that FetalAgents consistently delivers the most robust and accurate performance when compared against specialized models and multimodal large language models (MLLMs), ultimately providing an auditable, workflow-aligned solution for fetal ultrasound analysis and reporting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。