arXiv:2608.16033cs.CL2026-08

测试大模型在共享预算下的理性推理能力,发现其表现远低于潜力。

$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

论文配图:$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
图 1 · 摘自论文原文
  • 设计六题共用预算的评测基准,模拟真实资源约束场景。
  • 多数模型在共享预算下表现低于单题最优水平,差距达71个测试项。
  • 适合研究模型资源分配策略与认知理性,尤其关注智能体系统。

在认知科学中,资源理性探讨代理如何在计算资源有限时最大化预期收益。现有推理与智能体评测多采用独立任务预算;而现有共享预算研究未将套件性能与同一模型在单问题上的表现进行校准。我们提出 $R^3$-Bench,评估数学、编程竞赛与抽象推理三类任务,在无工具与代理设置下的六题共享预算表现。通过匹配单题响应曲线构建离线经验神谕,该神谕在72个主表单元中,平均表现均不低于或超过竞赛均值,且在71个单元中严格更高。在中等无工具压力下,均分重播策略也优于竞赛表现,覆盖六模型中的四个。轨迹诊断显示策略更新能力有限,失败模式受压力影响。在强代理压力下的三模型诊断中,至少一种固定调度器在九个单元中的六个超过竞赛均值,但无单一策略跨领域主导。结果揭示了模型已展示能力与共享预算实际表现之间的持续鸿沟。

原文摘要 · Abstract (English)

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.

资源理性智能体评测大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。