测试大模型在真实生活场景中的理解能力,发现现有模型表现仍很差。
CL-bench Life: Can Language Models Learn from Real-Life Context?

- 构建了405组真实生活上下文任务对,评估模型推理能力
- 顶尖模型仅19.3%任务解决率,平均仅13.8%
- 适合研究日常对话、个人行为记录等复杂场景的智能助手
当前的AI助手如OpenClaw需具备有效处理上下文的能力,而随着系统从专业场景进入日常生活,上下文也变得更加混乱、碎片化,并与个人及社会经验深度关联,如多方对话、个人档案和行为痕迹。然而,现有前沿语言模型是否能可靠地从这类上下文中学习并解决相关任务尚不明确。为此,我们提出了CL-bench Life,一个完全人工标注的基准,包含405个上下文-任务对和5,348条验证标准,覆盖常见真实生活场景。解决这些任务需要模型在复杂、混乱的真实上下文中进行推理,远超现有基准的评估范围。我们评估了十种前沿语言模型,发现真实生活上下文学习仍极具挑战:最优秀模型仅达19.3%任务解决率,平均为13.8%。模型在处理杂乱群聊历史或碎片化行为记录时仍表现不佳。该基准为推进真实生活上下文学习提供了关键测试平台,其进展将助力更智能可靠的日常智能助手发展。
原文摘要 · Abstract (English)
Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for models. As these systems move beyond professional settings into everyday life, the nature of the contexts they must handle also shifts. Real-life contexts are often messy, fragmented, and deeply tied to personal and social experience, such as multi-party conversations, personal archives, and behavioral traces. Yet it remains unclear whether current frontier language models can reliably learn from such contexts and solve tasks grounded in them. To this end, we introduce CL-bench Life, a fully human-curated benchmark comprising 405 context-task pairs and 5,348 verification rubrics, covering common real-life scenarios. Solving tasks in CL-bench Life requires models to reason over complex, messy real-life contexts, calling for strong real-life context learning abilities that go far beyond those evaluated in existing benchmarks. We evaluate ten frontier LMs and find that real-life context learning remains highly challenging: even the best-performing model achieves only 19.3% task solving rate, while the average performance across models is only 13.8%. Models still struggle to reason over contexts such as messy group chat histories and fragmented behavioral records from everyday life. CL-bench Life provides a crucial testbed for advancing real-life context learning, and progress on it can enable more intelligent and reliable AI assistants in everyday life.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。