评测大模型在真实智能家居中的推理与执行能力
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

- 基于可执行模拟器构建1100个分层任务
- 复杂家居环境下大模型自动化能力下降明显
- 适合研究智能助手、具身智能的开发者参考
智能家居正向状态依赖的复杂生活场景演进,要求大语言模型(LLMs)能够推理用户意图、偏好及多设备交互。然而现有基准多聚焦静态指令到API映射或有限仿真,无法评估大模型在真实家庭场景下的推理、交互与执行可靠性。为此,我们提出SMH-Bench,一个面向智能家居环境的综合性评估基准。基于可执行且可验证的HomeEnv模拟器,SMH-Bench包含1,100个高质量任务,覆盖7个类别和22个细粒度子类别,并按简单、中等、复杂三类住宅划分,涵盖从单间公寓到含135个设备的多房间环境。实验表明,尽管前沿大模型在明确控制和查询任务上表现良好,但在自动化任务调度、歧义处理和个性化推理方面仍存在显著不足,尤其在高复杂度环境中表现更差。我们期望SMH-Bench能推动更可靠、上下文感知且可实际部署的智能家居代理发展。
原文摘要 · Abstract (English)
Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferences, and multi-device interactions. However, existing smart-home benchmarks often focus on static instruction-to-API mapping or limited simulations, failing to evaluate whether LLMs can reason, interact, and act reliably in realistic household scenarios. To address these limitations, we introduce SMH-Bench, a comprehensive benchmark for evaluating LLMs in smart-home environments. Built upon HomeEnv, an executable and verifiable smart-home simulator, SMH-Bench contains 1,100 high-quality tasks spanning 7 categories and 22 fine-grained subcategories. It further stratifies tasks across simple, medium and complex homes, ranging from small apartments to dense multi-room environments with 135 devices. Experiments show that although frontier LLMs achieve strong performance on explicit control and query tasks, they still exhibit significant weaknesses in automation task scheduling, ambiguity handling and personalized reasoning, especially as home complexity increases. We hope SMH-Bench will facilitate the development of more reliable, context-aware, and practically deployable smart-home agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。