arXiv:2603.01357cs.AI2026-03被引 10

评测AI助手在个人上下文中的多步决策能力,发现现有模型在复杂场景下表现严重下降。

ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context

  • 构建动态个人上下文与工具交互结合的评测框架
  • 2413个场景测试显示高复杂度下模型性能显著下降
  • 适合研究具身智能、个性化助手和多步规划的学者使用

下一代AI需处理大量个人数据、多样工具和多步推理,但现有评测大多脱离上下文且为单轮。我们提出ASTRA-bench(助理技能在工具使用、推理与行动规划中的评测),首次将随时间演化的个人背景与交互式工具箱及复杂用户意图统一。通过事件驱动流程,生成了2413个基于长期生活事件的场景,并由人工标注其指代性、功能性和信息复杂度。对先进模型(如Claude-4.5-Opus、DeepSeek-V3.2)的评估显示,在高复杂度条件下性能明显下降,论证生成成为主要瓶颈。结果揭示当前智能体在混乱个人上下文中进行推理和可靠多步规划的能力存在重大局限。我们开源ASTRA-bench,包含完整执行环境与评估脚本,提供诊断性测试平台以推动真正上下文感知的AI助手发展。

原文摘要 · Abstract (English)

Next-generation AI must manage vast personal data, diverse tools, and multi-step reasoning, yet most benchmarks remain context-free and single-turn. We present ASTRA-bench (Assistant Skills in Tool-use, Reasoning \& Action-planning), a benchmark that uniquely unifies time-evolving personal context with an interactive toolbox and complex user intents. Our event-driven pipeline generates 2,413 scenarios across four protagonists, grounded in longitudinal life events and annotated by referential, functional, and informational complexity. Evaluation of state-of-the-art models (e.g., Claude-4.5-Opus, DeepSeek-V3.2) reveals significant performance degradation under high-complexity conditions, with argument generation emerging as the primary bottleneck. These findings expose critical limitations in current agents' ability to ground reasoning within messy personal context and orchestrate reliable multi-step plans. We release ASTRA-bench with a full execution environment and evaluation scripts to provide a diagnostic testbed for developing truly context-aware AI assistants.

智能代理多步规划上下文感知评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。