arXiv:2510.19771cs.AI2025-10被引 8

提出新评估框架,测试大模型主动解决问题能力

Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents

  • 将主动性拆解为搜寻问题、定位瓶颈、执行解决三步
  • 顶尖模型主动解决问题准确率仅达40%
  • 揭示大模型自主行动的局限与改进方向

基于大语言模型的智能体正迈向主动性:不再等待指令,而是主动预见用户需求并自主解决。然而,评估主动性仍具挑战性,现有基准局限于局部上下文,难以检验跨源推理与长时程规划能力。为此,我们提出PROBE(Proactive Resolution Of BottlEnecks)框架,将主动性分解为三个核心能力:(1)搜寻未明确指出的问题,(2)识别具体瓶颈,(3)执行相应解决方案。我们将PROBE应用于评估主流大模型与代理框架,发现即使是最先进的模型在该基准上也表现不佳。对前沿大模型与代理进行一致性测量后发现,最佳端到端性能为40%,由GPT-5与Claude Opus-4.1实现。此外,我们分析了各模型的相对能力并揭示其共性失败模式。结果凸显当前智能体自主行动的局限,并指明未来研究的重要方向。

原文摘要 · Abstract (English)

LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and solve them autonomously. However, evaluating proactivity is challenging; current benchmarks are constrained to localized context, limiting their ability to test reasoning across sources and longer time horizons. To address this gap, we present PROBE (Proactive Resolution Of BottlEnecks). PROBE decomposes proactivity as a pipeline of three core capabilities: (1) searching for unspecified issues, (2) identifying specific bottlenecks, and (3) executing appropriate resolutions. We apply PROBE to evaluate leading LLMs and popular agentic frameworks, showing that even state-of-the-art models struggle to solve this benchmark. Computing our consistent measurements across frontier LLMs and agents, we find that the best end-to-end performance of 40% is achieved by both GPT-5 and Claude Opus-4.1. Additionally, we demonstrate the relative capabilities of each model and analyze mutual failure modes. Our results highlight the current limitations of autonomous action in agentic systems, and expose promising future research directions.

大模型代理主动性评估智能体能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。