提出新基准与记忆压缩方法,让智能体更懂用户长期偏好。
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

- 设计多模态记忆压缩器,按用户偏好筛选关键信息
- 准确率提升7.18%,内存占用减少85.38%
- 适合研究长时决策与人机协同的学者
智能体需从长期交互中积累证据、理解隐含用户偏好,并在部分观察下比较多个选项。为此,本文提出DunphyBench基准,评估智能体在多场景住宅环境中完成以人为中心的长时决策任务的能力。不同于传统任务,该设置要求融合多源多模态输入以支持复杂推理。评估显示当前智能体与人类表现存在显著差距。诊断表明,原始多模态历史引入噪声,是性能瓶颈。为此,我们设计MeMento——一种基于用户偏好的多模态记忆压缩器,通过固定记忆令牌数量选择性压缩决策相关信息。实验表明,该方法使基于视觉语言模型的智能体准确率提升7.18%,内存使用降低85.38%。
原文摘要 · Abstract (English)
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over long horizons, interpret implicit user preferences, and compare multiple candidates under partial observations. In this work, we propose DunphyBench, a new benchmark for evaluating agents on long-horizon human-centered embodied decision-making, where the agent must navigate through multiple embodied housing environments and make decisions that align with multi-dimensional human preferences. Unlike standard embodied reasoning tasks that often focus on procedural planning or immediate goal completion, our setting requires agents to integrate multimodal, multi-source input into coherent knowledge that supports complex reasoning across long horizon. The evaluation results reveal that there is a substantial gap between current agents and human performance. Furthermore, our diagnosis of state-of-the-art VLM-driven agents reveals that memory management is one of the bottlenecks, where raw multimodal history introduces noise that hinders decision quality. Motivated by this finding, we design MeMento, a preference-conditioned multimodal memory compressor that selectively compresses decision-relevant information from long-horizon history based on user preferences with a fixed set of memory tokens. Experiments show that MeMento helps VLM-driven agents improve accuracy by 7.18%, while reducing memory usage by 85.38% compared to the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。