用大模型预测智能家居配置错误修复方案,效果显著。
Empirical evaluation of LLMs in predicting fixes of Configuration bugs in Smart Home System
- 对比GPT-4、GPT-4o和Claude 3.5 Sonnet在四种提示设计下的表现。
- 含原始代码的提示下,GPT-4与Claude 3.5 Sonnet策略预测准确率达80%。
- 提示设计影响大,完整信息提示最优,适合自动化运维研究者。
本实证研究评估了大型语言模型(LLMs)在预测智能家居系统配置错误修复方案中的有效性。研究分析了三种主流模型——GPT-4、GPT-4o(GPT-4 Turbo)和Claude 3.5 Sonnet——在四种不同提示设计下的表现,以评估其识别正确修复策略并生成有效解决方案的能力。研究使用来自Home Assistant Community的129个调试问题数据集,对其中21个随机选取案例进行了深入分析。结果表明,在提供错误描述和原始脚本的情况下,GPT-4与Claude 3.5 Sonnet在策略预测上分别达到80%的准确率。GPT-4在不同提示类型中表现稳定,而GPT-4o在速度和成本效益方面更具优势,尽管准确率略低。研究发现提示设计显著影响模型性能,包含描述与原始脚本的综合提示效果最佳。该研究为提升智能家居系统配置的自动修复能力提供了重要参考,并展示了LLMs在应对配置类问题方面的潜力。
原文摘要 · Abstract (English)
This empirical study evaluates the effectiveness of Large Language Models (LLMs) in predicting fixes for configuration bugs in smart home systems. The research analyzes three prominent LLMs - GPT-4, GPT-4o (GPT-4 Turbo), and Claude 3.5 Sonnet - using four distinct prompt designs to assess their ability to identify appropriate fix strategies and generate correct solutions. The study utilized a dataset of 129 debugging issues from the Home Assistant Community, focusing on 21 randomly selected cases for in-depth analysis. Results demonstrate that GPT-4 and Claude 3.5 Sonnet achieved 80\% accuracy in strategy prediction when provided with both bug descriptions and original scripts. GPT-4 exhibited consistent performance across different prompt types, while GPT-4o showed advantages in speed and cost-effectiveness despite slightly lower accuracy. The findings reveal that prompt design significantly impacts model performance, with comprehensive prompts containing both description and original script yielding the best results. This research provides valuable insights for improving automated bug fixing in smart home system configurations and demonstrates the potential of LLMs in addressing configuration-related challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。