让智能代理学会理解用户隐含偏好,自主完成个性化任务。
SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World
- 构建用户思维链,融合显性与隐性偏好进行推理
- 在新数据集SmartSpot上实现多任务个性化推荐
- 适合研究个性化智能体与人机交互的学者
基于大视觉语言模型的具身智能体在真实或虚拟世界中表现出色,但现有方法多依赖理想化动作轨迹和明确目标,忽视用户个体因素,导致在个性化服务中性能下降。为此,我们提出链式用户思维(Chain-of-User-Thought, COUT)框架,将用户行为从基础动作推理逐步过渡到显性和隐性偏好推理,融入个性化因素。基于此,我们构建SmartAgent框架,可感知虚拟环境、通过图形界面交互获取物品池信息、生成用户显性需求,并推荐满足其隐性需求的物品。为验证能力,我们创建了全新数据集SmartSpot,涵盖完整个性化交互场景。实验表明,SmartAgent在多项具身与个性化任务中表现优异。代码与数据将在论文录用后公开于https://github.com/tsinghua-fib-lab/SmartAgent。
原文摘要 · Abstract (English)
Recent advances in embodied agents with multimodal perception and reasoning capabilities based on large vision-language models (LVLMs), excel in autonomously interacting either real or cyber worlds, helping people make intelligent decisions in complex environments. However, the current works are normally optimized by golden action trajectories or ideal task-oriented solutions toward a definitive goal. This paradigm considers limited user-oriented factors, which could be the reason for their performance reduction in a wide range of personal assistant applications. To address this, we propose Chain-of-User-Thought (COUT), a novel embodied reasoning paradigm that takes a chain of thought from basic action thinking to explicit and implicit personalized preference thought to incorporate personalized factors into autonomous agent learning. To target COUT, we introduce SmartAgent, an agent framework perceiving cyber environments and reasoning personalized requirements as 1) interacting with GUI to access an item pool, 2) generating users' explicit requirements implied by previous actions, and 3) recommending items to fulfill users' implicit requirements. To demonstrate SmartAgent's capabilities, we also create a brand-new dataset SmartSpot that offers a full-stage personalized action-involved environment. To our best knowledge, our work is the first to formulate the COUT process, serving as a preliminary attempt towards embodied personalized agent learning. Our extensive experiments on SmartSpot illuminate SmartAgent's functionality among a series of embodied and personalized sub-tasks. We will release code and data upon paper notification at https://github.com/tsinghua-fib-lab/SmartAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。