arXiv:2606.26346cs.AI2026-06

评估工具增强大模型在真实能源市场分析任务中的表现

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

  • 用真实能源数据接口和专业工具评估大模型
  • 243个任务涵盖价格分析、收益建模等复杂场景
  • 开源数据集助力能源AI研究,适合从业者与研究者

当前代理评测多集中于通用或特定领域,但能源领域的评估仍以静态知识回忆为主,难以满足需实时数据获取、专业监管与市场知识及多步定量推理的现实需求。本文针对真实能源市场分析任务,开展工具增强型大模型的实证研究。构建包含243个专家标注问题的评估环境,覆盖三大类任务:(1)市场数据检索与分析,(2)知识检索与解读,(3)高级量化建模与决策分析。任务包括电价与需求分析、费率影响建模、资产收益估算、对冲策略分析与优化建模,难度分层。模型配备可配置领域工具,如美国主要电力市场ISO实时接口、监管文件检索、公用事业费率数据库、资产优化模型及能源文档RAG系统。采用多维评估协议,评分维度包括方法正确性、答案准确性、属性一致性与来源有效性,并按任务类型智能匹配评分标准。对比评估闭源与开源大模型,揭示模型能力与领域工具协同机制。关键数据集与工具链公开,支持复现与未来研究。

原文摘要 · Abstract (English)

Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a critical gap for a sector that requires live data retrieval, specialized regulatory and market knowledge, and multi-step quantitative reasoning under real-world constraints. We present an empirical study of tool-augmented LLM agents on real-world energy market analytics tasks. Our evaluation environment includes 243 expert-curated problems across three categories: (1) Market Data Retrieval and Analysis, (2) Knowledge Retrieval and Interpretation, and (3) Advanced Quantitative Modeling and Decision Analytics. Tasks include price and demand analysis, tariff impact modeling, asset revenue and returns estimation, hedging strategy analysis, and optimization modeling, with problems spanning multiple difficulty levels. Agents are equipped with a configurable suite of domain tools, including live electricity market APIs for major U.S. ISOs, regulatory docket search, utility tariff databases, asset optimization models, and retrieval-augmented generation over energy market documents. We assess agent responses using a multi-dimensional evaluation protocol that scores approach correctness, answer accuracy, attribute alignment, and source validity, with category-aware routing to match scoring criteria to question type. We evaluate both closed-source and open-source LLMs, providing a comparative analysis of how model capability and domain tooling interact in a high-stakes professional domain. Key artifacts are publicly released to support reproducibility and future research.

能源AI工具增强大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。