提出可保留气溶胶结构的低维嵌入方法,用于精准模拟气溶胶演化过程。
AeroMELD: A Linear Embedding of Aerosol Populations for Diagnostics and Latent Dynamics

- 基于尺度-形状分解构建线性编码器,保持气溶胶群体数学结构
- 在粒子解析数据上实现质量/数量分布、活化谱等高精度重建
- 适合需要物理可解释性的气候模型与混合机器学习研究者
准确表征大气气溶胶群体对模拟气溶胶-云相互作用、辐射强迫和冰核化至关重要,但现有简化方案依赖结构假设,难以捕捉组分多样性和混合状态。机器学习方法更具灵活性,但标准自编码器无法保留气溶胶群体的数学结构,因而不支持物理意义明确的过程算子。本文提出AeroMELD(气溶胶测度嵌入用于潜在动力学),一种数学严谨的低维潜在变量构建框架,能保持该结构。我们证明任意置换不变的线性编码器必须采用尺度-形状分解,总浓度显式表示,潜在形状由每颗粒子嵌入的重心组合构成。此聚合潜在状态通过将非线性后聚合阶段移入学习诊断映射,兼具Deep Sets模型的诊断表达力与潜在线性特性。使用粒子解析数据作为真实值,直接编码加权颗粒群体而非分箱气溶胶状态;尺寸分辨的质量与数量分布仅作为诊断目标和可视化摘要。潜在空间能精确重构这些分布、活化核谱、光学系数及浸没冻结行为,同时保持用于混合机器学习-物理模型所需的线性群体结构。尽管实验聚焦于诊断重建,该嵌入设计支持排放与混合的精确表示,并可在受控潜在空间中学习非线性微物理过程。本工作为直接在潜在空间学习气溶胶-过程演化奠定基础。
原文摘要 · Abstract (English)
Accurately representing atmospheric aerosol populations is essential for simulating aerosol-cloud interactions, radiative forcing, and ice nucleation, yet existing reduced schemes impose structural assumptions that limit their ability to capture composition diversity and mixing state. Machine-learning approaches offer more flexible representations, but standard autoencoders do not preserve the mathematical structure of aerosol populations and therefore cannot support physically meaningful process operators. We introduce AeroMELD (Aerosol Measure Embedding for Latent Dynamics), a mathematically grounded framework for constructing low-dimensional latent variables that retain this structure. We show that any permutation-invariant linear encoder must take a scale-shape decomposition, with total number concentration represented explicitly and latent shape given by a barycentric combination of per-particle embeddings. This aggregated latent state retains the diagnostic expressiveness of a Deep Sets model by moving the nonlinear post-aggregation stage into the learned diagnostic map while preserving latent linearity. Using particle-resolved data as ground truth, we encode weighted particle populations directly rather than binned aerosol states; size-resolved mass and number distributions serve only as diagnostic targets and visual summaries. The latent space accurately reconstructs these distributions, CCN spectra, optical coefficients, and immersion-freezing behavior while preserving the linear population structure needed for hybrid ML-physics models. Although the experiments focus on diagnostic reconstruction, the embedding is designed so that emissions and mixing can be represented exactly and nonlinear microphysical processes learned in a controlled latent space. This work establishes a foundation for learning aerosol-process evolution directly in latent space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。