用扩散模型检测科学AI预测的异常数据,给出可验证的信任证书。
Towards a Certificate of Trust: Task-Aware OOD Detection for Scientific AI
- 结合输入与预测结果,用扩散模型估计联合似然
- 在多个科学数据集上,似然值与误差强相关(R² > 0.8)
- 适合高风险科学场景中评估AI预测可信度
数据驱动模型在气象预报和流体动力学等关键科学领域应用日益广泛。这些方法在分布外(OOD)数据上可能失效,但回归任务中的失败检测仍是开放挑战。我们提出一种基于得分函数扩散模型估计联合似然的新方法,不仅考虑输入,还融合回归模型的预测结果,提供任务感知的可靠性评分。在包括偏微分方程(PDE)数据集、卫星图像和脑肿瘤分割在内的多个科学数据集上,该似然值与预测误差表现出强相关性(平均R² > 0.8)。本工作为构建可验证的‘信任证书’奠定了基础,为评估科学人工智能预测的可信度提供了实用工具。代码已公开:https://github.com/bogdanraonic3/OOD_Detection_ScientificML。
原文摘要 · Abstract (English)
Data-driven models are increasingly adopted in critical scientific fields like weather forecasting and fluid dynamics. These methods can fail on out-of-distribution (OOD) data, but detecting such failures in regression tasks is an open challenge. We propose a new OOD detection method based on estimating joint likelihoods using a score-based diffusion model. This approach considers not just the input but also the regression model's prediction, providing a task-aware reliability score. Across numerous scientific datasets, including PDE datasets, satellite imagery and brain tumor segmentation, we show that this likelihood strongly correlates with prediction error. Our work provides a foundational step towards building a verifiable 'certificate of trust', thereby offering a practical tool for assessing the trustworthiness of AI-based scientific predictions. Our code is publicly available at https://github.com/bogdanraonic3/OOD_Detection_ScientificML
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。