arXiv:2605.31281cs.CL2026-05

用大模型自动清理风电维护数据,提升故障分析准确性

Wind Turbine Maintenance Log Labelling Framework: LLM-Driven Data Correction and Enrichment via Semantic Extraction of Reliability Intelligence

  • 构建拓扑感知的LLM流程,从非结构化记录中提取故障与维修分类
  • 处理1.6万条记录,87%的标签达高置信度且无需人工复核
  • 输出带溯源的语义字段,适合做可靠性分析和故障预测研究

随着风力涡轮机队列老化,基于数据的可靠性工程与维护优化对控制全生命周期成本、延长资产寿命至关重要。历史维护记录是重要的现场证据来源,但其分析受制于不一致的系统编码、通用类别字段及非结构化技工文本。本文提出一种拓扑感知的大语言模型(LLM)工作流,用于审查旧标签、提取候选维修与故障模式分类体系,并在记录级别分配结构化语义字段。该工作流处理了来自32个陆上风电场、280台风机的16,316条维护记录,涵盖9.2年运行历史。流程结合系统特异性批处理与细粒度标注,采用确定性排除规则、结构化输出、记录级溯源及明确审查路径。针对三个系统编码任务的2,984条记录中,2,178条提出的标签满足‘高’自报告置信度标准且无需人工复核。维护类型与操作标签分别应用于14,251条和13,179条记录;故障模式证据档案分配至11,662条记录;3,441条因信息不足被标记,1,213条按确定性规则排除。结果揭示了系统与维护类型分布变化、更丰富的组件级操作词汇表,以及拓扑特定的候选证据档案。总API支出为368.86美元,即每条记录0.0226美元,总耗时6.83小时。带有溯源的输出可作为后续多源事件重构、基于暴露的可靠性分析及故障模式与影响分析(FMEA)的候选语义证据。

原文摘要 · Abstract (English)

As wind turbine fleets age, data-driven reliability engineering and maintenance optimisation are essential to manage lifecycle expenditure and support asset life extension. Historical maintenance records offer a vital source of field evidence, yet their analytical use is impeded by inconsistent system codes, generic categorical fields, and unstructured technician text. This paper presents a topology-aware large language model (LLM) workflow for reviewing legacy labels, extracting candidate maintenance and failure-mode taxonomies, and assigning structured semantic fields at record level. The workflow processed 16,316 maintenance records from 280 turbines across 32 onshore wind farms, spanning 9.2 years of operational history. It combines system-specific batch synthesis with granular labelling, deterministic exclusions, structured outputs, record-level provenance, and explicit review routes. Of 2,984 records targeted by three system-code tasks, 2,178 proposed labels met the operational acceptance rule of a 'High' self-reported confidence tier and no human-review flag. Accepted maintenance-type and action labels were assigned to 14,251 and 13,179 records, respectively. Failure-mode evidence profiles were assigned to 11,662 records; 3,441 records were classified as containing insufficient information, and 1,213 records were excluded as 'Not applicable' by deterministic workflow rules. The resulting fields reveal changes in system and maintenance-type distributions, a broader component-level action vocabulary, and topology-specific candidate evidence profiles. The recorded API expenditure was $368.86, or $0.0226 per processed record, and the total wall-clock duration was 6.83 hours. The provenance-linked outputs constitute candidate semantic evidence for subsequent multi-source event reconstruction, exposure-based reliability analysis, and failure modes and effects analysis (FMEA).

风电运维大模型应用数据清洗可靠性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。