arXiv:2607.11459eess.SYcs.AI2026-07

构建首个能源领域多模态数据集,支持大模型应用。

A Multimodal Dataset for Large Language Model Applications in the Energy Domain

论文配图:A Multimodal Dataset for Large Language Model Applications in the Energy Domain
图 1 · 摘自论文原文
  • 整合5万份文本、2万张图像等多源数据,统一格式。
  • 包含2500万条时间序列与200万条地理关系数据。
  • 符合FAIR原则,适合能源AI研究与决策系统开发。

本文提出mAIEnergy数据集,一个开放获取的多模态语料库,用于支持能源领域大语言模型(LLM)应用。该数据集整合了约5万份文本文档、2万张图像、2500万条数值时间序列记录以及200万条地理空间和关系型数据条目,涵盖政策法规、科研论文、新闻文章、卫星影像、电力系统测量、气象观测、统计指标及能源基础设施的空间表示。所有数据均已结构化处理,配备一致元数据和可复现的数据检索与预处理流程。该数据集可作为能源领域的基础知识库,供能源相关方集成开源或私有数据。mAIEnergy遵循可发现、可访问、可互操作、可重用(FAIR)原则,显著提升其在人工智能驱动的能源研究、建模与决策中的应用价值。

原文摘要 · Abstract (English)

This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observations, statistical indicators, and geospatial representations of energy infrastructure and related entities. All data have been harmonized into structured, ready-to-use formats, accompanied by consistent metadata and reproducible data retrieval and preparation workflows. The dataset can serve as a foundational energy knowledge base, allowing energy stakeholders to integrate additional open-source or proprietary data. The mAIEnergy dataset adheres to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, enhancing its applicability for AI-driven energy research, modeling, and decision-making.

多模态数据能源AI大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。