让机器学习数据自带因果机制和不确定性,提升科学模型可信度。
Instrumented data for causal scientific machine learning
- 数据附带生成它的机制模型与不确定性,支持反事实推演
- 实现可验证的图像到模拟的全流程推理,参数可编辑可追溯
- 适合需要可解释性与因果干预的科研场景,如生物、气候建模
科学机器学习的瓶颈不在于模型规模,而在于数据。观测数据只记录结果,不说明原因;模板合成数据虽有明确生成过程,但仅适用于模拟器预设模板,无法覆盖用户实际问题。我们提出第三种可行方案:仪器化数据,即每个数据点都包含其生成机制模型、对模型的显式不确定性估计,以及可执行的反事实族。验证与验证(V&V)框架下的仪器化图像-模拟管道是具体实现:传感器观测转化为完整可追溯、求解器支撑的模拟,具备可编辑参数及传播的随机性/认知不确定性。该数据基座具有案例特异性、机制监督性,并可通过Pearl的do算子支持因果干预。短期内可应用于计算生物学、气候、材料、流体力学与医学影像的验证、审计与代理训练;长期来看,可能催生可被证伪的科学推理基础模型。
原文摘要 · Abstract (English)
Scientific machine learning is limited less by model size than by the data it is trained on. Observational data records what happened but not why; template synthetic data has a known generating process but only for the simulator's template, not the case a user faces. We argue a third option is now operationally feasible: instrumented data, in which every datum carries the mechanistic model that produced it, an explicit uncertainty over that model, and an executable family of counterfactuals. Verification-and-validation (V&V) instrumented image-to-simulation pipelines are one realisation: a sensor observation becomes a fully specified, solver-backed simulation with explicit, editable parameters and a propagated aleatoric/epistemic uncertainty. The substrate is case-specific, mechanistically supervised, and supports causal interventions through Pearl's do-operator. Near-term consequences for validation, auditing, and surrogate training span computational biology, climate, materials, fluid mechanics, and medical imaging; a longer-term, falsifiable implication concerns foundation models for scientific reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。