用可验证仿真构建智能家庭代理训练数据飞轮,提升任务成功率。
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation

- 通过程序化生成家居环境与任务目标,实现多样化训练数据自动生成。
- 在SmartHome-Bench上达成87.03%任务成功率,超过GPT-5.5 1.23个百分点。
- 适合研究智能体训练、物理世界交互与基于反馈的强化学习方向者。
大型语言模型代理正从纯文本交互迈向真实世界控制,智能家庭是代表性领域。真实家庭交互需理解模糊意图、适应动态环境并完成多轮推理,但现有方法难以生成高质量训练数据。本文提出HomeFlow,一种可验证的数据飞轮机制。该系统采用HomeEnv作为统一仿真环境,HomeMaker用于程序化生成多样家居场景;Blueprint将开放式用户意图转化为可执行的状态成功条件,MCTS-Flow则通过环境引导的树搜索合成多样且可验证的多轮操作轨迹。随后,通过监督微调与逐步强化验证(step-wise RLVE)优化代理,实现基于真实物理反馈的迭代改进。我们进一步构建了SmartHome-Bench基准评估代理性能。在该基准上,HomeFlow-RL-4B和HomeFlow-RL-8B的任务成功率达84.60%和87.03%,其中HomeFlow-RL-8B甚至超越领先模型GPT-5.5达1.23个百分点。
原文摘要 · Abstract (English)
Large language model agents are moving beyond text-only interaction toward physical-world control, with smart homes as a representative domain. Real domestic interaction requires understanding ambiguous intents, operating in dynamic environments, and performing multi-turn reasoning. However, existing methods struggle to generate high-quality training data for smart home agents. We propose HomeFlow, a verifiable data flywheel for this domain. HomeFlow uses HomeEnv as a unified simulation environment and HomeMaker to procedurally generate diverse home settings. Subsequently, Blueprint compiles open-ended user intents into executable state-based success conditions, while MCTS-Flow synthesizes diverse, verifiable multi-turn trajectories through environment-guided tree search. We then optimize the agents via supervised fine-tuning and step-wise RLVE, which facilitates iterative improvement through authentic physical feedback. We further construct SmartHome-Bench to evaluate the agent across various smart home tasks. On this benchmark, HomeFlow-RL-4B and HomeFlow-RL-8B achieve task success rates of 84.60% and 87.03%. It is worth noting that HomeFlow-RL-8B even surpasses the leading GPT-5.5 by 1.23 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。