arXiv:2603.06636cs.LGcs.AI2026-03被引 2

首个面向大模型的智能家居异常检测数据集,评测其在真实场景下的故障感知能力。

SmartBench: Evaluating LLMs in Smart Homes with Anomalous Device States and Behavioral Contexts

  • 构建包含正常与异常设备状态及行为上下文的智能家庭数据集
  • 13个主流大模型在无上下文异常检测中最高仅66.1%准确率
  • 揭示当前大模型在智能家居异常识别上仍有显著短板,适合智能系统研究者

由于大语言模型(LLMs)展现出强大的上下文感知能力,近期研究开始探索将其集成到智能家居助手以帮助用户管理生活环境。尽管已有研究表明大模型能有效理解用户需求并给出合适响应,但现有研究多集中于解释和执行用户指令,而忽视了智能助手关键功能——检测环境异常状态的能力。这要求大模型不仅能准确判断是否存在异常,还需提供清晰解释或可操作建议。为提升下一代基于大模型的智能家居助手的异常检测能力,我们提出SmartBench,这是首个专为大模型设计的智能家居数据集,包含正常与异常设备状态以及正常与异常的状态转移上下文。我们在该基准上评估了13个主流大模型。实验结果表明,大多数先进模型在异常检测上表现不佳:例如Claude-Sonnet-4.5在无上下文异常类别上仅达到66.1%的检测准确率,在依赖上下文的异常上更低至57.8%。更多实验表明,下一代基于大模型的智能家居助手仍远未具备有效检测和应对家庭异常状况的能力。我们的数据集已公开于https://github.com/horizonsinzqs/SmartBench。

原文摘要 · Abstract (English)

Due to the strong context-awareness capabilities demonstrated by large language models (LLMs), recent research has begun exploring their integration into smart home assistants to help users manage and adjust their living environments. While LLMs have been shown to effectively understand user needs and provide appropriate responses, most existing studies primarily focus on interpreting and executing user behaviors or instructions. However, a critical function of smart home assistants is the ability to detect when the home environment is in an anomalous state. This involves two key requirements: the LLM must accurately determine whether an anomalous condition is present, and provide either a clear explanation or actionable suggestions. To enhance the anomaly detection capabilities of next-generation LLM-based smart home assistants, we introduce SmartBench, which is the first smart home dataset designed for LLMs, containing both normal and anomalous device states as well as normal and anomalous device state transition contexts. We evaluate 13 mainstream LLMs on this benchmark. The experimental results show that most state-of-the-art models cannot achieve good anomaly detection performance. For example, Claude-Sonnet-4.5 achieves only 66.1% detection accuracy on context-independent anomaly categories, and performs even worse on context-dependent anomalies, with an accuracy of only 57.8%. More experimental results suggest that next-generation LLM-based smart home assistants are still far from being able to effectively detect and handle anomalous conditions in the smart home environment. Our dataset is publicly available at https://github.com/horizonsinzqs/SmartBench.

大模型智能家居异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。