提出可适配数据量与任务类型的超声多任务学习框架,避免统一训练性能下降。
Understanding Task Aggregation for Generalizable Ultrasound Foundation Models
- 基于DINOv3设计任务自适应的Mixture-of-Experts结构,动态分配模型容量。
- 27个超声任务实验证明:数据少时联合训练易产生负迁移,数据多时临床分组更优。
- 揭示分割任务对聚合最敏感,指导未来模型设计应结合任务特性与数据规模。
基础模型有望在单一框架内统一多个临床任务,但近期超声研究显示,统一模型可能劣于专用基线。我们假设性能下降并非由模型容量限制导致,而是任务聚合策略忽视了任务异质性与可用训练数据规模之间的相互作用。本文系统分析了异质性超声任务在何种条件下可无损联合学习,为统一临床影像模型的任务聚合提供实用准则。提出M2DINO,一个基于DINOv3的多器官、多任务框架,采用任务条件化的Mixture-of-Experts模块实现容量自适应分配。我们在三种范式下评估了27个超声任务(涵盖分割、分类、检测、回归),结果表明聚合效果强烈依赖训练数据规模:在数据充足时,临床分组训练可提升性能;但在低数据场景下会引发显著负迁移。相比之下,全任务统一训练在各临床组间表现更一致。实验还发现任务敏感度随任务类型变化:分割任务性能下降最明显,优于回归与分类任务。这些发现为超声基础模型设计提供实践指导,强调聚合策略应同时考虑训练数据可用性与任务特征,而非仅依赖临床分类体系。
原文摘要 · Abstract (English)
Foundation models promise to unify multiple clinical tasks within a single framework, but recent ultrasound studies report that unified models can underperform task-specific baselines. We hypothesize that this degradation arises not from model capacity limitations, but from task aggregation strategies that ignore interactions between task heterogeneity and available training data scale. In this work, we systematically analyze when heterogeneous ultrasound tasks can be jointly learned without performance loss, establishing practical criteria for task aggregation in unified clinical imaging models. We introduce M2DINO, a multi-organ, multi-task framework built on DINOv3 with task-conditioned Mixture-of-Experts blocks for adaptive capacity allocation. We systematically evaluate 27 ultrasound tasks spanning segmentation, classification, detection, and regression under three paradigms: task-specific, clinically-grouped, and all-task unified training. Our results show that aggregation effectiveness depends strongly on training data scale. While clinically-grouped training can improve performance in data-rich settings, it may induce substantial negative transfer in low-data settings. In contrast, all-task unified training exhibits more consistent performance across clinical groups. We further observe that task sensitivity varies by task type in our experiments: segmentation shows the largest performance drops compared with regression and classification. These findings provide practical guidance for ultrasound foundation models, emphasizing that aggregation strategies should jointly consider training data availability and task characteristics rather than relying on clinical taxonomy alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。