arXiv:2608.10875cs.CLcs.AI2026-08

评测大模型在长期生活任务中的主动性和持续性表现。

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

论文配图:VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
图 1 · 摘自论文原文
  • 设计200个跨10个生活场景的长期任务,模拟真实世界变化。
  • 所有模型在新基准上得分均低,暴露当前智能体能力短板。
  • 适合研究长期决策、自主行为的AI系统开发者使用。

大型语言模型代理正被广泛用作个人助手,但现有评估多基于静态环境中的短时、独立请求。真实生活辅助则不同:任务持续数周而非几分钟,世界持续变化而代理未被主动提示,许多约束也未明说。仅回应眼前问题的代理在此类任务中必然失败。真正需要的是能主动决策、保持一致的代理——自行判断何时行动、何时提问、何时沉默;察觉未被宣布的变化;从第一天到最后一日维持计划连贯性。现有基准无法衡量此类能力。我们提出VibeLifeBench,包含200个长期任务,覆盖十类日常生活领域,运行于由22个模拟服务构成的动态世界中。世界按自身时钟推进,许多变化无声发生,唯有主动重查才能发现。每项任务通过细粒度加权检查评分,仅依据代理实际留下的行为,涵盖最终状态、行动及时性及隐含约束遵守情况。我们评估了七种前沿模型,结果均表现不佳,揭示当前代理与真实生活协助能力之间的巨大差距。我们将开源全部任务、环境与评估框架。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.

智能体长期任务主动行为评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。