arXiv:2606.30647cs.CVcs.AI2026-06

融合多源观测数据,实现云微物理场高精度四维重建。

Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation

论文配图:Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation
图 1 · 摘自论文原文
  • 通过跨模态层级注意力融合图像与雷达/测云仪数据
  • 液态水含量误差0.026 g/m³,风速误差1.18 m/s
  • 适合气象建模与智能观测系统研究者

从稀疏的地基观测仪器中进行云微物理场的密集体积重建仍是开放问题,主要因测量数据在模态和空间覆盖上差异大。我们提出AtmoFuseNet框架,融合多视角天空相机图像、毫米波云雷达和测云仪观测,生成云状态与风场的4D(三维空间加时间)估计。方法分三阶段:第一阶段为跨模态层级聚合模块,通过逐层交叉注意力将图像特征金字塔与仪器反演的垂直剖面结合;第二阶段为条件变分精修模块,在可微分雷达与图像前向模型约束下,将结果体数据映射为物理一致的微物理场;第三阶段为基于相关性的运动估计算法,从连续体重建中恢复每一体素的3D风矢量。在半干旱站点的共址观测数据上,AtmoFuseNet达到0.026 g/m³液态水含量平均绝对误差和1.18 m/s风速平均绝对误差,优于现有反演基准。消融实验量化了各模块贡献。

原文摘要 · Abstract (English)

Dense volumetric reconstruction of cloud microphysical fields from sparse ground-based instruments remains an open problem, largely because the available measurements are heterogeneous in both modality and spatial coverage. We present AtmoFuseNet, a framework that fuses multi-view sky camera imagery with millimeter-wave cloud radar and ceilometer observations to produce 4D (three spatial dimensions plus time) estimates of cloud state and wind. The method operates in three stages: a cross-modal hierarchical aggregation module that combines image feature pyramids with instrument-derived vertical profiles through layer-wise cross-attention; a conditional variational refinement module that maps the resulting volume to physically consistent microphysical fields under differentiable radar and image forward models; and a correlation-based motion estimator that recovers per-voxel 3D wind vectors from consecutive volumetric reconstructions. On collocated observations from a semi-arid site, AtmoFuseNet reaches 0.026 g m^-3 liquid water content MAE and 1.18 m s^-1 wind speed MAE, improving over existing retrieval baselines. Ablation experiments isolate the contribution of each module.

云重建多模态融合4D气象传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。