arXiv:2608.28063cs.CV2026-08

对比复杂与简单模型,发现简单方法在超声多器官分类中表现更优且更可靠。

A Controlled Audit of Architectural Complexity in Uncertainty-Aware Multi-Organ Ultrasound Classification

论文配图:A Controlled Audit of Architectural Complexity in Uncertainty-Aware Multi-Organ Ultrasound Classification
图 1 · 摘自论文原文
  • 设计受控实验框架,系统评估复杂模型与简化版本的性能差异。
  • 简单交叉熵+温度缩放在双数据集上均达到最优校准效果。
  • 复杂组件如不确定性门控和证据学习在实际中作用有限,需实证支持留存。

多器官超声分类器日益融合注意力机制、专家混合路由、不确定性门控和证据深度学习(EDL)目标以应对解剖异质性和采集差异。然而,设计合理性并不等同于性能提升。本文提出受控复杂性审计框架,用于评估最大证据模型Full-EDL与简化替代方案的部署决策。在主数据集及三个内部复现实验中,使用十组随机种子、冻结图像级划分、容量与优化感知比较、对称温度缩放、成对决策规则及独立的分布外(OOD)否决机制进行评估。结果显示,保留Full-EDL在任一数据集上均未带来可靠的宏平均F1提升,而简化方案在非劣边界下仍不明确。简单交叉熵+温度缩放(Simple-CE+TS)在两个数据集上均满足校准负对数似然标准,并表现出有利的选择性风险排序。原始证据训练的校准优势在温度缩放后消失,且在第二数据集未重现。门控模块在审计检查点处影响可忽略,删除仅全模型链也未带来稳定任务或校准损失收益。尽管Simple-CE触发了胎儿探头的OOD否决,但未触发肺部探头,因此无法做出无条件的OOD安全性声明。故选定Simple-CE+TS作为评估的分布内目标,同时保留Full-EDL作为最大参考。组件应通过功能与重训练证据证明其留存价值,校准与分布偏移可靠性需分开评估。

原文摘要 · Abstract (English)

Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition. Yet a plausible design rationale does not by itself establish that an added component improves the trained system. We contribute a controlled complexity-audit framework, applied to the deployment decision between the maximal evidential candidate Full-EDL and simpler alternatives. Six candidates were evaluated on the primary dataset and three in an internal replication, using ten matched seeds, frozen image-level partitions, capacity- and optimisation-aware comparisons, symmetric temperature scaling, paired decision rules, and a separate out-of-distribution (OOD) veto. Retaining Full-EDL did not establish a reliable macro-F1 gain on either dataset, while the simplified alternatives remained inconclusive under the non-inferiority margin. Simple cross-entropy with temperature scaling (Simple-CE+TS) met the calibrated negative log-likelihood criterion on both datasets and showed favourable selective-risk ordering. The raw calibration advantage of evidential training disappeared after temperature scaling and did not recur on the second dataset. The gate had negligible observable influence at the audited checkpoints, and deleting the Full-only chain revealed no stable task or calibrated-loss benefit. Simple-CE nevertheless triggered the OOD veto against the fetal probe but not the lung probe, precluding an unconditional OOD-safety claim. We therefore selected Simple-CE+TS for the evaluated in-distribution objective while retaining Full-EDL as the maximal reference. Components should earn retention through functional and retraining-based evidence, and calibration and distribution-shift reliability should be evaluated separately.

医学影像模型复杂度不确定性建模超声分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。