arXiv:2604.16813cs.AIcs.CL2026-04被引 1

评测智能家庭中AI代理的个性化推理与规划能力。

PersonalHomeBench: Evaluating Agents in Personalized Smart Homes

论文配图:PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
图 1 · 摘自论文原文
  • 构建动态家庭状态生成个性化任务,模拟真实家居环境。
  • 复杂任务下模型表现下降,尤其在部分观测和反事实推理中失败。
  • 适合研究个性化AI代理、智能家庭系统及多模态交互的学者。

智能体式AI系统正快速迈向现实应用,但在复杂个性化环境中的准备程度仍不充分。为填补这一空白,我们提出PersonalHomeBench,一个用于评估基础模型在个性化智能家居环境中作为智能助手表现的基准。该基准通过迭代过程逐步构建丰富的家庭状态,并据此生成个性化、上下文相关的任务。为支持真实的智能体-环境交互,我们提供PersonalHomeTools工具箱,涵盖家庭信息检索、设备控制与情境理解功能。PersonalHomeBench在单模态与多模态观测下评估智能体的反应式与主动性能力。大量实验表明,随着任务复杂度提升,模型性能系统性下降,尤其在反事实推理和部分可观测条件下出现明显失败,此时需依赖工具进行有效信息获取。这些结果使PersonalHomeBench成为分析个性化智能体推理与规划鲁棒性及局限性的严格评估平台。

原文摘要 · Abstract (English)

Agentic AI systems are rapidly advancing toward real-world applications, yet their readiness in complex and personalized environments remains insufficiently characterized. To address this gap, we introduce PersonalHomeBench, a benchmark for evaluating foundation models as agentic assistants in personalized smart home environments. The benchmark is constructed through an iterative process that progressively builds rich household states, which are then used to generate personalized, context-dependent tasks. To support realistic agent-environment interaction, we provide PersonalHomeTools, a comprehensive toolbox enabling household information retrieval, appliance control, and situational understanding. PersonalHomeBench evaluates both reactive and proactive agentic abilities under unimodal and multimodal observations. Thorough experimentation reveals a systematic performance reduction as task complexity increases, with pronounced failures in counterfactual reasoning and under partial observability, where effective tool-based information gathering is required. These results position PersonalHomeBench as a rigorous evaluation platform for analyzing the robustness and limitations of personalized agentic reasoning and planning.

智能家庭智能体评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。