构建首个智能家居记忆控制评估基准,提升长期记忆管理能力
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards

- 提出多维奖励强化学习框架,实现对设备操作的细粒度记忆控制
- 发布真实用户日志构建的MemHomeLife数据集,支持长期行为分析
- 设计首个系统性评估记忆驱动控制的基准MemHome,适配家庭场景研究者
大型语言模型已成为实现个性化智能家居体验的关键基础。尽管现有研究探索了智能家居助手如何实时理解用户指令并控制设备,但在基于记忆的设备控制方面仍面临评估与方法论双重挑战。评估层面,现有基准或聚焦即时设备控制,或关注通用开放域记忆检索任务,无法有效评估模型在记忆驱动控制中的表现。方法层面,虽可用强化学习实现记忆驱动控制,但传统强化学习依赖结果反馈(即任务是否完成),缺乏中间反馈,导致在细粒度记忆管理任务(如添加、更新、删除、使用)中性能不佳或出现局部失败。为此,我们首次发布基于匿名真实长期用户交互日志的MemHomeLife数据集;为进一步实现对不同记忆相关子任务的精细评估,我们构建了首个专为智能家庭场景设计的系统性评估基准MemHome。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become a key foundation for enabling personalized smart home experiences. While existing studies have explored how smart home assistants understand user queries to control devices in real time, their ability to perform memory-driven device control remains challenging from both evaluation and methodological perspectives. In terms of evaluation, existing benchmarks either focus on immediate device control or general open-domain memory retrieval tasks, and therefore cannot effectively evaluate a model's ability to perform memory-driven device control. Methodologically, while memory-driven device control can be approached using Reinforcement Learning, conventional RL methods generally rely on outcome-based supervision (i.e., whether the final task is achieved). This lack of intermediate feedback can lead to sub-optimal performance or local failures in fine-grained memory management tasks (adding, updating, deleting, and utilizing). To address these issues, we first release MemHomeLife, built from anonymized real-world long-term user interaction logs. To enable more fine-grained evaluation of different memory-related subtasks, we further construct MemHome, the first benchmark designed to systematically evaluate memory-driven device control in smart home scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。