arXiv:2509.25748cs.CVcs.AI2025-09

首个统一多任务的超声影像大模型,提升诊断准确率与可解释性。

Dolphin v1.0 Technical Report

  • 构建200万样本多模态数据集,融合真实、合成与知识库数据
  • 在U2-Bench上达0.5835分,性能超第二名近一倍
  • 引入强化学习增强推理能力,适合医疗AI研发与临床辅助场景

超声在现代医学中至关重要,但存在操作者依赖、图像噪声和实时扫描难题,阻碍AI应用。尽管大型多模态模型在其他医学影像领域表现优异,却难以应对超声的复杂性。为此,我们提出Dolphin v1.0(V1)及其推理增强版Dolphin R1——首个大规模多模态超声基础模型,统一多种临床任务于单一视觉-语言框架。为应对超声数据变异性和噪声问题,我们构建了200万规模的多模态数据集,涵盖教材知识、公开数据、合成样本及通用语料,确保模型具备鲁棒感知、泛化与临床适应能力。Dolphin系列采用三阶段训练策略:领域特化预训练、指令驱动对齐、强化学习优化。Dolphin v1.0在分类、检测、回归和报告生成任务中表现可靠;Dolphin R1通过超声专用奖励机制的强化学习,提升诊断推理、透明度与可解释性。在U2-Bench上的八项任务评估中,Dolphin R1取得0.5835的U2得分,超过第二名模型(0.2968)近一倍,创下新纪录。Dolphin v1.0也表现优异,验证了统一框架的有效性。对比表明,推理增强训练显著提升诊断准确性、一致性与可解释性,凸显其在高风险医疗AI中的关键价值。

原文摘要 · Abstract (English)

Ultrasound is crucial in modern medicine but faces challenges like operator dependence, image noise, and real-time scanning, hindering AI integration. While large multimodal models excel in other medical imaging areas, they struggle with ultrasound's complexities. To address this, we introduce Dolphin v1.0 (V1) and its reasoning-augmented version, Dolphin R1-the first large-scale multimodal ultrasound foundation models unifying diverse clinical tasks in a single vision-language framework.To tackle ultrasound variability and noise, we curated a 2-million-scale multimodal dataset, combining textbook knowledge, public data, synthetic samples, and general corpora. This ensures robust perception, generalization, and clinical adaptability.The Dolphin series employs a three-stage training strategy: domain-specialized pretraining, instruction-driven alignment, and reinforcement-based refinement. Dolphin v1.0 delivers reliable performance in classification, detection, regression, and report generation. Dolphin R1 enhances diagnostic inference, reasoning transparency, and interpretability through reinforcement learning with ultrasound-specific rewards.Evaluated on U2-Bench across eight ultrasound tasks, Dolphin R1 achieves a U2-score of 0.5835-over twice the second-best model (0.2968) setting a new state of the art. Dolphin v1.0 also performs competitively, validating the unified framework. Comparisons show reasoning-enhanced training significantly improves diagnostic accuracy, consistency, and interpretability, highlighting its importance for high-stakes medical AI.

超声影像多模态模型医疗AI推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。