arXiv:2605.24558cs.LG2026-05

AI for Science应把测量到数据的流程当作可计算的推断组件,而非固定数据接口。

Position: AI for Science Should Treat Measurement-to-Dataset Pipelines as Inference Components

论文配图:Position: AI for Science Should Treat Measurement-to-Dataset Pipelines as Inference Components
图 1 · 摘自论文原文
  • 将测量到数据的多阶段流程视为可推断的组件,而非固定输入。
  • 跨数据集稳定性测试中仅0.0004%的流程能存活,暴露隐性假设与不确定性。
  • 适合关注可复现、可审计科学证据的研究者,尤其在神经科学等间接观测领域。

AI for Science工作流常将发布数据集视为系统固定接口。但在依赖间接观测的领域,学习者实际观察的是经多阶段测量、重建和预处理流程生成的衍生表示。我们主张:这些从测量到数据的流程本质上是推断组件。若将其输出当作‘给定数据’,则会冻结观测模型,掩盖对可行流程选择的不确定性。我们识别出三种由此导致的失效模式:(C1) 隐性假设空间——数据集未说明流程配置及其有效条件;(C2) 未经认证的可迁移性——流程虽有文档但有效性范围未验证,分布偏移下的失败无法判定;(C3) 无管控的多重性——存在多种合理流程,其差异真实存在却未被纳入不确定性传播。通过大规模神经科学实证审计,我们发现跨数据集稳定性标准下,仅有约0.0004%的流程能存活。我们呼吁AI4Science社区通过特定领域的可计算观测框架,使流程成为可计算的推断对象。此举可量化流程合理性与稳定性,将隐含实现选择转化为可审计、可复现、可累积的科学证据。

原文摘要 · Abstract (English)

AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on \emph{indirect observation}, the learner observes a derivative representation produced by multi-stage measurement, reconstruction, and preprocessing pipelines. \textbf{We argue that these measurement-to-dataset pipelines are inference components: treating their outputs as ``given data'' freezes an observation model and obscures uncertainty over feasible pipeline choices.} We identify three failure modes arising from this ``frozen lens'': \textbf{(C1) hidden hypothesis space}, where the released dataset does not specify the pipeline configuration or its validity conditions; \textbf{(C2) uncertified transportability}, where a pipeline may be documented but its regime of validity is untested, so failures under distribution shift cannot be adjudicated; \textbf{(C3) ungoverned multiplicity}, where many defensible pipelines exist and dispersion is real but not propagated into uncertainty-aware evidence. We stress-test these claims with a large-scale neuroscience empirical audit, finding a survival rate of $\approx 0.0004\%$ under a cross-dataset stability criterion. We call on the AI4Science community to make pipelines \emph{computable} inference objects via domain-specific Computable Observation Frameworks. This shift enables quantifying pipeline adequacy and stability, converting implicit implementation choices into auditable, reproducible, and cumulative scientific evidence.

AI for Science可计算观测不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。