arXiv:2605.26256cs.AI2026-05

让智能体通过长期互动记忆,理解用户隐含需求。

Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions

论文配图:Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
图 1 · 摘自论文原文
  • 构建多模态知识图谱,存储用户交互的语义与视觉记忆
  • 在跨轮次任务中提升性能,多跳推理准确率显著提高
  • 适合需要长期个性化服务的智能助手场景

基于多模态大语言模型(MLLM)的具身智能体在物理环境中展现出解决复杂任务的强大潜力。然而,个性化服务不仅依赖通用指令或物体识别,更需理解用户通过长期互动积累的隐含意图。本文提出POLAR,一种基于多模态记忆增强的具身智能体框架,将过往交互组织为包含语义记忆(如个人偏好)和情景记忆(如智能体轨迹)的多模态知识图谱。执行任务时,系统检索相关记忆以解析当前请求并指导行动。我们在多个MLLM骨干网络及多样评估场景下验证了该框架,结果表明:记忆机制能持续提升性能,尤其在跨多轮推理、多跳推断或追踪用户上下文变化时效果显著。

原文摘要 · Abstract (English)

Multimodal large language model (MLLM)-based embodied agents have shown strong potential for solving complex tasks in physical environments. However, personalized assistance requires more than following generic instruction or recognizing object categories. In real-world scenarios, the intended target is often specified only implicitly through prior interactions, requiring agents to leverage personalized context accumulated over time. In this work, we propose POLAR, a multiomodal memory-augmented framework for personalized embodied agents over long-term user interactions. POLAR organizes prior interactions into a multimodal knowledge graph that captures semantic memory for personalized context and visual concepts, and episodic memory for embodied experiences such as agent trajectories. To execute embodied tasks, POLAR retrieves relevant memories to interpret the current request and guide task execution. We evaluate POLAR across multiple MLLM backbones and diverse evaluation scenarios to study the role of memory in long-term personalization. Results show that the proposed memory mechanism consistently improves performance by enabling more effective use of information accumulated over prior interactions. The gains are especially pronounced when the agents are required to reason across multiple interactions, perform multi-hop inference, or tracking updates in user-specific context over time.

具身智能多模态记忆个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。