arXiv:2606.02863cs.AI2026-06被引 1

提出GAMBLe框架,解析AI科研系统中组件如何影响搜索效果。

Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems

论文配图:Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
图 1 · 摘自论文原文
  • 将AI科研系统拆解为生成器、评估器等四参数与有效景观
  • 实验显示组件组合差异可提升性能67%、效率39倍
  • 适合关注AI辅助科研的算法设计者和系统优化研究者

AI驱动的研究系统(ADRS)通过大模型与自动化评估结合发现算法、证明和设计,正被广泛采用,但其分析工具滞后。ADRS性能依赖于组件间复杂互动,这些机制难以理解且成本高昂,标准收敛性保证在此过程中不适用。本文提出GAMBLe框架,将ADRS行为分解为生成器$G$、评估器$A$、发现机制$M$、预算$B$四个参数及复合对象有效景观$L_{\text{eff}} = \mathcal{A} \circ G$,揭示不同生成器-评估器组合导致结构各异的优化景观。在760+次重复实验(超46,000次迭代)中,覆盖单个大模型至动态自适应集成生成器、贪心选择至协同进化元搜索机制,以及三个NP难问题,评估方式从连续评分到悬崖函数。结果表明无整体最优生成器或机制:前沿模型可能逊于开源替代品,最简机制有时优于先进元搜索。即使在每轮仅60次迭代的有限预算下,正确组件组合仍可使性能提升13%-67%,搜索效率提高6-39倍。

原文摘要 · Abstract (English)

AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace. ADRS performance depends on component interactions that are poorly understood, expensive to explore, and (as we show) not well captured by standard convergence guarantees. These guarantees rely on structural assumptions that do not hold under the ADRS process we formalize. We introduce GAMBLe, a framework that decomposes ADRS behavior into four parameters (generator $G$, assessor $\mathcal{A}$, discovery mechanism $\mathcal{M}$, budget $B$) and one compositional object, the effective landscape $L_{\text{eff}} = \mathcal{A} \circ G$, which reveals that distinct generator-assessor pairs induce structurally different per-problem optimization landscapes. We exercise the framework on 760+ replicated runs (>46,000 iterations) spanning generators from single LLMs to dynamically-adaptive ensembles, mechanisms from greedy selection to co-evolutionary meta-search, and three NP-hard problems whose assessors range from continuous scoring to cliff functions. The experiments reveal no total ordering of generators or mechanisms: frontier models can underperform open-source alternatives and the simplest mechanism sometimes outperforms state-of-the-art meta-search. Results show that even under limited budgets (60 iterations per run), the right component choices can improve performance by 13-67% and search efficiency by 6-39x.

AI科研大模型优化框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。