评测大模型对风机运维日志的分类能力,发现最佳模型可高精度识别部件但难判维修意图。
A Comparative Benchmark of Large Language Models for Labelling Wind Turbine Maintenance Logs
- 构建开源框架,系统评估多款大模型在风机日志分类任务中的表现。
- 顶级模型在部件识别上准确率高,但对维修动作解释的一致性较差。
- 建议采用人机协作模式,提升运维数据标注效率与质量。
风力发电的高效运行与维护(O&M)对降低度电成本(LCOE)至关重要,但风机运维日志多为非结构化自由文本,阻碍了自动化分析。本文提出一个新颖且可复现的大语言模型(LLM)基准测试框架,用于评估其在复杂工业记录分类任务中的表现。该框架已开源,以促进透明性和后续研究。我们系统评估了多种先进的专有及开源LLM,全面分析其在可靠性、运行效率和模型校准方面的权衡。结果量化出明确的性能层级,识别出在基准标准上高度对齐且置信度可信的领先模型。同时发现,分类性能显著受任务语义模糊性影响:所有模型在客观部件识别上共识更高,而在解释性维修动作上表现差异大。由于无模型达到完美准确率且校准差异显著,我们结论是近中期最有效且负责任的应用方式是人机协同系统——让LLM作为强大助手,加速并标准化人类专家的数据标注,从而提升运维数据质量与下游可靠性分析能力。
原文摘要 · Abstract (English)
Effective Operation and Maintenance (O&M) is critical to reducing the Levelised Cost of Energy (LCOE) from wind power, yet the unstructured, free-text nature of turbine maintenance logs presents a significant barrier to automated analysis. Our paper addresses this by presenting a novel and reproducible framework for benchmarking Large Language Models (LLMs) on the task of classifying these complex industrial records. To promote transparency and encourage further research, this framework has been made publicly available as an open-source tool. We systematically evaluate a diverse suite of state-of-the-art proprietary and open-source LLMs, providing a foundational assessment of their trade-offs in reliability, operational efficiency, and model calibration. Our results quantify a clear performance hierarchy, identifying top models that exhibit high alignment with a benchmark standard and trustworthy, well-calibrated confidence scores. We also demonstrate that classification performance is highly dependent on the task's semantic ambiguity, with all models showing higher consensus on objective component identification than on interpretive maintenance actions. Given that no model achieves perfect accuracy and that calibration varies dramatically, we conclude that the most effective and responsible near-term application is a Human-in-the-Loop system, where LLMs act as a powerful assistant to accelerate and standardise data labelling for human experts, thereby enhancing O&M data quality and downstream reliability analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。