构建可执行的车载多用户长期记忆评测基准,检验智能助手在动态偏好下的持续适应能力。
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
- 基于可执行仿真环境,通过对比动作后状态与目标状态评估记忆与工具使用。
- 每样本含80+历史记忆事件,23个工具模块,揭示大模型在动态偏好下表现不佳。
- 适合研究车载智能体长期记忆、多用户冲突处理与自适应决策的学者和工程师。
随着对智能车载体验需求的增长,车载代理正从简单助手演变为长期陪伴者。这要求代理能持续建模多用户偏好,并在用户间偏好冲突及习惯变化下做出可靠决策。然而,现有基准大多局限于单用户、静态问答场景,无法捕捉偏好随时间演变及真实车载环境中多用户、工具交互的特性。为此,我们提出VehicleMemBench,一个基于可执行车载仿真环境的多用户长上下文记忆评测基准。该基准通过比较动作后环境状态与预设目标状态,实现无需大模型或人工评分的客观可复现评估。基准包含23个工具模块,每个样本包含超过80条历史记忆事件。实验表明,尽管强大模型在直接指令任务中表现良好,但在涉及记忆演化的场景中表现不佳,尤其在用户偏好动态变化时。即使先进记忆系统也难以应对该环境中的领域特定记忆需求。这些发现凸显了构建更鲁棒、专用记忆管理机制以支持真实车载系统长期自适应决策的必要性。为促进未来研究,我们公开数据与代码。
原文摘要 · Abstract (English)
With the growing demand for intelligent in-vehicle experiences, vehicle-based agents are evolving from simple assistants to long-term companions. This evolution requires agents to continuously model multi-user preferences and make reliable decisions in the face of inter-user preference conflicts and changing habits over time. However, existing benchmarks are largely limited to single-user, static question-answer settings, failing to capture the temporal evolution of preferences and the multi-user, tool-interactive nature of real vehicle environments. To address this gap, we introduce VehicleMemBench, a multi-user long-context memory benchmark built on an executable in-vehicle simulation environment. The benchmark evaluates tool use and memory by comparing the post-action environment state with a predefined target state, enabling objective and reproducible evaluation without LLM-based or human scoring. VehicleMemBench includes 23 tool modules, and each sample contains over 80 historical memory events. Experiments show that powerful models perform well on direct instruction tasks but struggle in scenarios involving memory evolution, particularly when user preferences change dynamically. Even advanced memory systems struggle to handle domain-specific memory requirements in this environment. These findings highlight the need for more robust and specialized memory management mechanisms to support long-term adaptive decision-making in real-world in-vehicle systems. To facilitate future research, we release the data and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。