arXiv:2606.03331cs.CLcs.AI2026-06

评测大模型在真实设备维修中的表现,发现其存在安全风险且跨语言能力弱。

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

论文配图:Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions
图 1 · 摘自论文原文
  • 构建991个真实维修问题数据集,含技术员参考答案,覆盖手机、电脑和数据恢复。
  • 所有模型在主板诊断和安全操作上错误率高,GPT-5.4表现最优。
  • 跨语言性能差,孟加拉语回答普遍不如英语,高危任务需严格安全防护。

消费设备维修是大型语言模型(LLMs)的一个重要但未被充分探索的测试场景。维修任务需要对不完整的问题描述进行推理,结合硬件特异性诊断、可操作的排错步骤以及关乎安全的关键决策,错误建议可能导致设备损坏、电池危险或永久性数据丢失。我们引入一个包含991个来自Reddit的真实维修问题的数据集,涵盖手机维修、计算机维修和数据恢复,每个问题均配有技师撰写的参考解决方案,并提供孟加拉语翻译以评估跨语言性能。我们使用四项维修特定标准——正确性、完整性、实用性和安全性——在英语和孟加拉语中评估六种最先进的LLM。结果表明,尽管LLMs能提供有用的维修协助,但在缺乏严格评估和明确安全防护的情况下,仍不可靠地应用于高风险实际维修任务。手机维修是最具挑战性和安全敏感的领域,所有模型在板级诊断、维修优先级排序和安全恢复流程上均出现显著错误。跨领域与模型分析显示,孟加拉语响应表现持续劣于英语。在所评估模型中,GPT-5.4整体表现最佳。

原文摘要 · Abstract (English)

Consumer device repair is an important but underexplored testbed for large language models (LLMs). Repair tasks require reasoning over incomplete problem descriptions, hardware-specific diagnostics, actionable troubleshooting, and safety-critical decisions, where incorrect advice can cause device damage, battery hazards, or permanent data loss. We introduce a benchmark of 991 real-world repair questions from Reddit spanning phone repair, computer repair, and data recovery, each paired with technician-written reference solutions, and provide Bangla translations to evaluate cross-lingual performance. We evaluate six state-of-the-art LLMs in English and Bangla using four repair-specific criteria: correctness, completeness, practicality, and safety. Our results show that while LLMs can provide useful repair assistance, they remain unreliable for high-risk real-world repair tasks without rigorous evaluation and explicit safety safeguards. Phone repair is the most difficult and safety-sensitive domain, and all models make substantial errors in board-level diagnosis, repair prioritization, and safe recovery procedures. Across domains and models, Bangla responses consistently perform worse than English responses. Among the evaluated models, GPT-5.4 performs best overall.

大模型评测设备维修安全风险多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。