厘清机器学习不确定性量化中意图与实现的错位问题
On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning
- 区分不同场景下的不确定性类型与建模目标
- 指出当前方法在频率派、贝叶斯等框架间映射存在错误
- 适合关注可信机器学习系统的研究人员参考
机器学习模型的不确定性量化是现代数据分析的核心挑战。该挑战源于两个关键方面:(a) 不同学科对不确定性及其估计术语使用不一致,(b) 多样化问题背景下建立可信不确定性所需的技术要求各异。本文通过识别这些不一致之处,阐明不同应用场景带来的认知需求差异。我们分析了当前的估计目标(如预测、推断、基于模拟的推断)、不确定性构造(如频率主义、贝叶斯、似然性)及其相互映射方式。结合文献,揭示并解释了若干不当映射的典型案例。为应对这些问题,我们倡导建立标准,以促进不确定性量化方法中‘意图’与‘实现’的对齐。讨论了确保可靠不确定性的多个可信维度,并说明其如何指导不确定性感知机器学习系统的设计与评估。实践建议聚焦科学机器学习,提供典型示例与使用场景,尤其在基于模拟的推断(SBI)背景下。
原文摘要 · Abstract (English)
Quantifying uncertainties for machine learning (ML) models is a foundational challenge in modern data analysis. This challenge is compounded by at least two key aspects of the field: (a) inconsistent terminology surrounding uncertainty and estimation across disciplines, and (b) the varying technical requirements for establishing trustworthy uncertainties in diverse problem contexts. In this position paper, we aim to clarify the depth of these challenges by identifying these inconsistencies and articulating how different contexts impose distinct epistemic demands. We examine the current landscape of estimation targets (e.g., prediction, inference, simulation-based inference), uncertainty constructs (e.g., frequentist, Bayesian, fiducial), and the approaches used to map between them. Drawing on the literature, we highlight and explain examples of problematic mappings. To help address these issues, we advocate for standards that promote alignment between the \textit{intent} and \textit{implementation} of uncertainty quantification (UQ) approaches. We discuss several axes of trustworthiness that are necessary (if not sufficient) for reliable UQ in ML models, and show how these axes can inform the design and evaluation of uncertainty-aware ML systems. Our practical recommendations focus on scientific ML, offering illustrative cases and use scenarios, particularly in the context of simulation-based inference (SBI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。