arXiv:2510.12953cs.CVcs.AI2025-10KDD被引 2

FetalMind让AI更懂胎儿超声,提升报告生成与诊断准确率。

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

  • 通过临床流程引导的解耦机制,分离视角与疾病关联
  • 在12家医院数据上训练,关键病种准确率提升61.2%
  • 适合产科医生辅助诊断,尤其对复杂病例有帮助

近期医疗视觉-语言模型在VQA、报告生成和异常检测任务中表现良好,但大多针对成人结构化影像,在胎儿超声任务中表现不佳,面临多视角推理、疾病种类繁多及图像多样性等挑战。为此,我们提出专用于胎儿超声的FetalMind系统,支持报告生成与诊断。基于临床工作流程,设计了显著认知解耦(SED)机制,将专家标注的二分图注入模型,通过强化学习解耦视图-疾病关联并引导临床合理的偏好选择,降低疾病间差异与视图异质性带来的学习瓶颈,使模型推理更贴合产科实践。为规模化训练,我们构建了首个大规模胎儿超声报告语料库FetalSigma-1M,包含来自12家医疗机构的20,000份报告,缓解领域数据稀缺问题。大量实验表明,FetalMind在所有孕周阶段均优于开源与闭源基线,平均性能提升14%,关键病症准确率提高61.2%,同时保持高效、稳定与可扩展性。

原文摘要 · Abstract (English)

Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which poses challenges of multi-view image reasoning, numerous diseases, and image diversity. To bridge this gap, we introduce FetalMind, a medical AI system tailored to fetal ultrasound for both report generation and diagnosis. Guided by clinical workflow, we propose Salient Epistemic Disentanglement (SED), which injects an expert-curated bipartite graph into the model to decouple view-disease associations and to steer preference selection along clinically faithful steps via reinforcement learning. This design mitigates variability across diseases and heterogeneity across views, reducing learning bottlenecks while aligning the model's inference with obstetric practice. To train FetalMind at scale, we curate FetalSigma-1M dataset, the first large-scale fetal ultrasound report corpus, comprising 20K reports from twelve medical centers, addressing the scarcity of domain data. Extensive experiments show that FetalMind outperforms open- and closed-source baselines across all gestational stages, achieving +14% average gains and +61.2% higher accuracy on critical conditions while remaining efficient, stable, and scalable. Project Page: https://hexiao0275.github.io/FetalMind.

胎儿超声视觉语言模型医学AI报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。