用可视化工具让气象数据嵌入表示更科学可解释,支持交互式气候发现。
Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration

- 构建视觉分析工作台,连接嵌入实验与原始数据、地理信息和模型配置。
- 在数十亿嵌入中实现交互式搜索,响应时间低于2秒,可在普通硬件上运行。
- 适合气候科学家探索极端天气事件,尤其适用于基于深度学习的气象数据挖掘。
地球系统科学正产生大量高维数据,来自物理模型与人工智能驱动模型。尽管嵌入表示使数据可检索,并成为智能发现引擎的基础,但潜在空间中的最近邻未必具有科学意义,可能反映预处理、地理分布或模型偏差。研究人员亟需可视化工具来检验潜空间结构,追溯检索结果的物理依据,并对比不同表示。本文提出一个开源的视觉分析工作台,支持可溯源的科学检索流程。系统将不同嵌入实验与原始数据、元数据、空间上下文及模型配置关联。用户可进行图像级与局部块级查询,施加多约束过滤,并通过熟悉的气象视图检查相似案例。这形成了科学家先在已知数据中识别现象,再用其潜空间特征探测更大数据集的发现闭环。我们以DINOv3在ERA5数据上检索热带气旋为例展示,该框架对模型无依赖,未来可集成其他嵌入架构。最后评估了其外存检索后端,证明在普通硬件上对数千万嵌入实现低延迟交互式搜索具有高度可扩展性。
原文摘要 · Abstract (English)
Earth system science is producing increasingly large, high-dimensional datasets from both physics-based and AI-driven models. While embedding-based representations make these data searchable and serve as foundational building blocks for AI-driven discovery engines, nearest neighbors in latent spaces are not automatically scientifically meaningful. They may reflect real meteorological structures, or simply artifacts of preprocessing, geography, or model bias. Researchers therefore need visual tools to inspect latent space organization, trace search results back to physical evidence, and evaluate candidate representations against one another. We present an open source visual analytics workbench designed to support this provenance-aware scientific retrieval workflow. The system links distinct embedding experiments to shared source data, metadata, spatial contexts, and model configurations. It enables interactive retrieval strategy design by allowing users to issue image-level and localized patch-level queries, apply multi-constraint filters, and inspect analogs through familiar meteorological views. This facilitates a discovery loop where scientists characterize a phenomenon in a well-understood dataset and use its latent signature to probe larger archives. While we demonstrate the workbench through a tropical cyclone retrieval scenario using a vision foundation model (DINOv3) on ERA5 data, the framework is model-agnostic and designed to integrate with other embedding architectures in the future. Finally, we evaluate its out-of-core retrieval backend, demonstrating that interactive visual search over tens of millions of embeddings is highly scalable on commodity hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。