arXiv:2410.09084cs.CLcs.AI2024-10被引 1

用大模型诊断机器人系统故障,效果媲美GPT-4且更便宜

Diagnosing Robotics Systems Issues with Large Language Models

  • 用微调的大模型分析机器人故障日志,找根因
  • 70亿参数模型经QLoRA微调后准确率超GPT-4
  • 适合工业机器人运维人员快速定位问题

在工业应用中快速解决故障对减少经济损失至关重要。然而,即使对专家而言,分析数据以确定根本原因仍是一项耗时且复杂的任务。相比之下,大语言模型(LLMs)擅长处理大量数据。已有研究证明其在IT系统运维中的有效性。本文将该方法拓展至尚待探索的机器人系统领域。我们构建了SYSDIAGBENCH——一个包含2500多个报告故障的专有机器人系统诊断基准。基于此,我们评估了不同规模模型与适配技术在根因分析中的表现。结果表明,采用QLoRA微调的70亿参数模型在诊断准确率上可超越GPT-4,同时成本显著更低。通过人工专家验证,最佳模型的判断结果与参考标签一致。

原文摘要 · Abstract (English)

Quickly resolving issues reported in industrial applications is crucial to minimize economic impact. However, the required data analysis makes diagnosing the underlying root causes a challenging and time-consuming task, even for experts. In contrast, large language models (LLMs) excel at analyzing large amounts of data. Indeed, prior work in AI-Ops demonstrates their effectiveness in analyzing IT systems. Here, we extend this work to the challenging and largely unexplored domain of robotics systems. To this end, we create SYSDIAGBENCH, a proprietary system diagnostics benchmark for robotics, containing over 2500 reported issues. We leverage SYSDIAGBENCH to investigate the performance of LLMs for root cause analysis, considering a range of model sizes and adaptation techniques. Our results show that QLoRA finetuning can be sufficient to let a 7B-parameter model outperform GPT-4 in terms of diagnostic accuracy while being significantly more cost-effective. We validate our LLM-as-a-judge results with a human expert study and find that our best model achieves similar approval ratings as our reference labels.

机器人大模型故障诊断运维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。