arXiv:2507.19408cs.LGcs.AI2025-07

同一患者不同模型可能给出冲突诊断,小集成可显著降低这种不确定性。

On Arbitrary Predictions from Equally Valid Models

  • 用多个等效模型对比预测,发现单模型选择存在任意性。
  • 小集成加弃权策略能有效减少预测分歧,高一致性的结果可自动分类。
  • 模型能力越强,预测分歧越小,但准确率本身无法根治该问题。

模型多重性指多个机器学习模型对数据描述能力相当,却在个别样本上产生不同预测。在医学领域,这可能导致同一患者获得不一致诊断,风险未被充分认识和应对。本研究在多种医疗任务和模型架构下,实证分析了预测多重性的程度、驱动因素及后果,发现(1)标准验证指标无法识别唯一最优模型;(2)大量预测依赖于模型开发过程中的任意选择。使用多个等效模型可揭示预测差异——若仅用单个模型,部分患者将面临随意诊断。相反,(3)小规模集成结合弃权策略可在实践中有效缓解可度量的预测多重性;高模型间一致性预测可适于自动化分类。尽管准确率本身并非解决多重性的根本方法,我们发现(4)通过增加模型容量提升准确率可减少预测多重性。研究强调需重视临床中的模型多重性,建议采用集成策略提升诊断可靠性;当模型无法达成足够共识时,应转交专家审定。

原文摘要 · Abstract (English)

Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflicting predictions for the same patient -- a risk that is poorly understood and insufficiently addressed. In this study, we empirically analyze the extent, drivers, and ramifications of predictive multiplicity across diverse medical tasks and model architectures, and show that even small ensembles can mitigate/eliminate predictive multiplicity in practice. Our analysis reveals that (1) standard validation metrics fail to identify a uniquely optimal model and (2) a substantial amount of predictions hinges on arbitrary choices made during model development. Using multiple models instead of a single model reveals instances where predictions differ across equally plausible models -- highlighting patients that would receive arbitrary diagnoses if any single model were used. In contrast, (3) a small ensemble paired with an abstention strategy can effectively mitigate measurable predictive multiplicity in practice; predictions with high inter-model consensus may thus be amenable to automated classification. While accuracy is not a principled antidote to predictive multiplicity, we find that (4) higher accuracy achieved through increased model capacity reduces predictive multiplicity. Our findings underscore the clinical importance of accounting for model multiplicity and advocate for ensemble-based strategies to improve diagnostic reliability. In cases where models fail to reach sufficient consensus, we recommend deferring decisions to expert review.

模型多重性医疗诊断集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。