arXiv:2604.04202cs.LGcs.AI2026-04被引 8

测试AI助手在不断变化的信息环境中保持正确认知的能力。

ClawArena: Benchmarking AI Agents in Evolving Information Environments

论文配图:ClawArena: Benchmarking AI Agents in Evolving Information Environments
图 1 · 摘自论文原文
  • 设计多源冲突与动态信念更新的评估场景
  • 模型表现差14.5分,框架设计差24分,表明系统架构影响更大
  • 适合研究智能体长期推理与个人化学习的学者

部署为持续助理的AI智能体必须随信息环境演化而更新信念。现实中,证据分散于异构来源,常相互矛盾,新信息可能推翻旧结论,用户偏好通过修正而非明确指令呈现。现有基准大多假设静态、单一权威设定,无法评估智能体应对复杂性的能力。我们提出ClawArena,一个面向动态信息环境的智能体评估基准。每个场景维持完整隐藏真实状态,仅向智能体暴露跨渠道会话、工作区文件和分阶段更新中的噪声、部分且时常矛盾的信息流。评估围绕三大耦合挑战:多源冲突推理、动态信念更新与隐式个性化展开,形成14类问题分类体系。采用选择题与可执行命令检查两种格式,分别检验推理与工作区对齐能力。ClawArena包含12个多轮场景,共337次评估回合与45次动态更新,覆盖五种智能体框架及18种语言模型(来自专有、社区可访问与自托管来源)。实验显示,模型能力导致29分得分差异,框架设计影响达24分;MetaClaw技能叠加能稳定提分而不降准确率;信念更新难度由更新策略决定,而非更新数量。代码已开源:https://github.com/aiming-lab/ClawArena。

原文摘要 · Abstract (English)

AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is scattered across heterogeneous sources that often contradict one another, new information can invalidate earlier conclusions, and user preferences surface through corrections rather than explicit instructions. Existing benchmarks largely assume static, single-authority settings and do not evaluate whether agents can keep up with this complexity. We introduce ClawArena, a benchmark for evaluating AI agents in evolving information environments. Each scenario maintains a complete hidden ground truth while exposing the agent only to noisy, partial, and sometimes contradictory traces across multi-channel sessions, workspace files, and staged updates. Evaluation is organized around three coupled challenges: multi-source conflict reasoning, dynamic belief revision, and implicit personalization, whose interactions yield a 14-category question taxonomy. Two question formats, multi-choice (set-selection) and shell-based executable checks, test both reasoning and workspace grounding. ClawArena comprises 12 multi-turn scenarios spanning 337 evaluation rounds with 45 dynamic updates, evaluated across five agent frameworks and 18 language models from proprietary, community-accessible, and self-hosted sources. Experiments show that model capability accounts for a 29-point score range across models while framework design accounts for up to a 24-point range, that MetaClaw's skill overlay reliably improves score without degrading accuracy, and that belief revision difficulty is determined by update design strategy rather than update volume. Code is available at https://github.com/aiming-lab/ClawArena.

智能体评估信念更新多源推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。