用分布式生成模型统一分析多源异构数据,提升反问题求解精度。
Multi-Dataset Inverse Problem Solving with Distributed Generative AI

- 将生成式反问题求解扩展至非同分布数据,共享参数约束全局一致
- 在模拟散射实验中验证,对不同探测器精度具有鲁棒性
- 适合实验条件不一致的多源数据联合分析,如高能物理、材料科学
从多个异构数据集中提取一组未知且不可直接测量的量,是科学领域的常见挑战。例如,来自不同探测器分辨率的测量数据。若仅独立或简单合并分析,难以获得精确无偏估计;需妥善处理数据异质性,但计算成本高。本文提出一种通用框架,用于在基于生成式AI的反问题求解中同时分析多个异构数据集。基于近期提出的可扩展异步生成反问题求解器(SAGIPS),我们将分布式数据并行训练范式扩展至非同分布数据:各数据集由相同未知推断参数控制,但覆盖特征空间不同区域。每个数据集通过独立前向算子与判别器处理,提供互补约束,共同引导共享生成器实现全局参数一致性。我们在受控的多探测器散射实验设置下验证该方法。数值结果表明,该框架对因未知探测器系统误差引起的多种数据保真度具有鲁棒性,并展示了在多GPU超算系统上的良好扩展性。结果证明该方法适用于实验条件各异的真实多数据集分析。
原文摘要 · Abstract (English)
Extracting a shared set of unknown, not directly measurable quantities from multiple, heterogeneous datasets is a common challenge across scientific domains. A prominent example is the combination of datasets obtained from different measurements with different settings (e.g. varying detector resolutions). Analyzing such datasets jointly, rather than independently or after naive merging, is essential for obtaining precise and unbiased estimates of the unknowns, but requires careful treatment of dataset heterogeneity and is computationally demanding. We present a generalized framework for simultaneously analyzing multiple heterogeneous datasets in the context of generative AI-based inverse problem solvers. Building on our recent Scalable Asynchronous Generative Inverse Problem Solver (SAGIPS) framework, we extend the well-established distributed data-parallel training paradigm to non-identically distributed datasets, where each dataset is controlled by the same set of unknown inference parameters but covers a different region of the available feature space. Each dataset is processed through its own forward operator and discriminator, providing complementary constraints that collectively guide a shared generator toward global parameter consistency. We validate the approach using a controlled setup inspired by a multi-detector scattering experiment. We provide numerical evidence that our framework is robust to different data fidelities, which arise from unknown detector systematics in the Rutherford experiment, and we show the scaling behavior on multi-GPU leadership computing systems. The results show that our approach is well suited for real-world multi-dataset analyses in which experimental conditions vary across measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。