arXiv:2507.22911cs.CLcs.AI2025-07

构建电力营销领域大模型评估基准,提升对话合规性与稳定性。

ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing

  • 设计跨6大服务域24子场景的55万条对话数据集,融合人工评分与合规测试。
  • 7B模型经优化后性能超更大模型,计算成本更低且符合监管要求。
  • 适合电力系统、智能客服研发人员及政策合规研究者参考。

随着电力系统低碳化与数字化发展,分布式能源高渗透率和灵活电价使电力营销(EPM)成为监管、系统运行与可持续能源部署的关键接口。当前许多电力公司仍依赖人工客服或基于规则/意图的聊天机器人,知识库碎片化,难以应对长周期、跨场景对话,无法满足合规、可验证及需求响应(DR)就绪的要求。与此同时,前沿大语言模型(LLMs)虽具强对话能力,但现有通用评测基准低估了行业术语、监管推理与多轮对话稳定性。为此,我们提出ElectriQ——一个面向电力营销的大型基准与评估框架。ElectriQ包含超过55万条对话,覆盖六个服务领域和24个子场景,并定义统一协议,结合人工评分、自动指标及两项合规压力测试:法定引用准确性(Statutory Citation Correctness)与长对话一致性(Long-Dialogue Consistency)。基于ElectriQ,我们提出SEEK-RAG,一种在微调与推理阶段注入政策与领域知识的检索增强方法。在13个LLM上的实验表明,经领域对齐的7B模型配合SEEK-RAG,性能可媲美甚至超越更大模型,同时显著降低计算开销,为支持需求侧管理、可再生能源集成与电网韧性运行的可审计、监管感知型电力营销助手提供了坚实基础。

原文摘要 · Abstract (English)

As power systems decarbonise and digitalise, high penetrations of distributed energy resources and flexible tariffs make electric power marketing (EPM) a key interface between regulation, system operation and sustainable-energy deployment. Many utilities still rely on human agents and rule- or intent-based chatbots with fragmented knowledge bases that struggle with long, cross-scenario dialogues and fall short of requirements for compliant, verifiable and DR-ready interactions. Meanwhile, frontier large language models (LLMs) show strong conversational ability but are evaluated on generic benchmarks that underweight sector-specific terminology, regulatory reasoning and multi-turn process stability. To address this gap, we present ElectriQ, a large-scale benchmark and evaluation framework for LLMs in EPM. ElectriQ contains over 550k dialogues across six service domains and 24 sub-scenarios and defines a unified protocol that combines human ratings, automatic metrics and two compliance stress tests-Statutory Citation Correctness and Long-Dialogue Consistency. Building on ElectriQ, we propose SEEK-RAG, a retrieval-augmented method that injects policy and domain knowledge during finetuning and inference. Experiments on 13 LLMs show that domain-aligned 7B models with SEEK-RAG match or surpass much larger models while reducing computational cost, providing an auditable, regulation-aware basis for deploying LLM-based EPM assistants that support demand-side management, renewable integration and resilient grid operation.

电力营销大模型评测合规推理RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。