SimuHome打造动态智能家庭环境,评估大模型代理的长期交互能力。
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- 基于Matter协议构建动态家庭模拟器,支持设备操作对温湿度等环境变量的实时影响
- 涵盖600个任务周期,工作流调度失败率最高,跨框架仍难解决
- 适合研究智能体长期决策、环境感知与真实世界前验验证的团队
我们提出SimuHome,一个高保真智能家庭模拟器及包含600个任务周期的基准测试。现有基准将家庭视为静态系统,既不模拟设备操作对环境变量(如温度、湿度)的时序影响,也不支持设备命令的工作流调度。SimuHome基于行业标准Matter协议构建,代理通过API与设备交互,并持续观测动作对环境变量的影响。基准覆盖状态查询、隐式用户意图推断、显式设备控制和工作流调度四类任务,每类均包含可行与不可行请求。针对工作流调度,模拟器可加速时间以实现即时评估。对18个代理的评测显示,工作流调度是难度最高的任务类别,失败现象在不同代理框架与微调策略中普遍存在。结果表明,SimuHome的时间加速模拟可作为代理在真实世界执行前的预验证环境。
原文摘要 · Abstract (English)
We introduce $\textbf{SimuHome}$, a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents. Existing smart home benchmarks treat the home as a static system, neither simulating how device operations affect environmental variables over time nor supporting workflow scheduling of device commands. SimuHome is grounded in the Matter protocol, the industry standard that defines how real smart home devices communicate and operate. Agents interact with devices through SimuHome's APIs and observe how their actions continuously affect environmental variables such as temperature and humidity. Our benchmark covers state inquiry, implicit user intent inference, explicit device control, and workflow scheduling, each with both feasible and infeasible requests. For workflow scheduling, the simulator accelerates time so that scheduled workflows can be evaluated immediately. An evaluation of 18 agents reveals that workflow scheduling is the hardest category, with failures persisting across alternative agent frameworks and fine-tuning. These findings suggest that SimuHome's time-accelerated simulation could serve as an environment for agents to pre-validate their actions before committing them to the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。