分离多机制数据中的独立影响,提升科学数据建模的可解释性。
DIVIDE: A Framework for Learning from Independent Multi-Mechanism Data Using Deep Encoders and Gaussian Processes
- 用深度编码器与高斯过程联合建模,分离空间、类别等独立机制
- 在合成数据和实验数据上准确还原加性与缩放交互关系
- 适合需要可解释性的科学建模,尤其适用于多物理场数据
科学数据常源于多个独立机制(如空间、类别或结构效应),其共同作用会掩盖各因素的独立贡献。我们提出DIVIDE框架,通过机制特异性深度编码器与结构化高斯过程在联合隐空间中联合建模,实现对独立生成因子的解耦。编码器分离不同机制,高斯过程捕捉其联合效应并提供校准的不确定性估计。该架构支持结构化先验,实现可解释的机制感知预测与高效主动学习。在合成数据(含类别图像块与非线性空间场)、FerroSIM铁电晶格模拟及PbTiO3薄膜的实验性PFM滞回环数据上验证,DIVIDE能有效分离机制、复现加性与缩放交互,并在噪声下保持鲁棒性。该框架自然拓展至机械、电磁或光学响应共存的多功能数据集。
原文摘要 · Abstract (English)
Scientific datasets often arise from multiple independent mechanisms such as spatial, categorical or structural effects, whose combined influence obscures their individual contributions. We introduce DIVIDE, a framework that disentangles these influences by integrating mechanism-specific deep encoders with a structured Gaussian Process in a joint latent space. Disentanglement here refers to separating independently acting generative factors. The encoders isolate distinct mechanisms while the Gaussian Process captures their combined effect with calibrated uncertainty. The architecture supports structured priors, enabling interpretable and mechanism-aware prediction as well as efficient active learning. DIVIDE is demonstrated on synthetic datasets combining categorical image patches with nonlinear spatial fields, on FerroSIM spin lattice simulations of ferroelectric patterns, and on experimental PFM hysteresis loops from PbTiO3 films. Across benchmarks, DIVIDE separates mechanisms, reproduces additive and scaled interactions, and remains robust under noise. The framework extends naturally to multifunctional datasets where mechanical, electromagnetic or optical responses coexist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。