arXiv:2512.16733cs.AI2025-12

用随机搜索主动测试AI能力,快速搞清它能做什么、何时能做。

Monte Carlo Query Search: Active Capability Assessment of AI Agents

  • 通过蒙特卡洛树搜索生成高区分度的测试问题
  • 在多个AI系统上比基线方法更快学出准确能力模型
  • 适合想评估复杂AI系统可靠性的研究者与工程师

黑箱AI系统(如基础模型代理)越来越多地用于序列决策。安全部署需要明确其能力边界、触发条件及可能结果。我们提出蒙特卡洛查询合成(MCQS),一种主动查询生成方法,用于学习黑箱AI的符号化随机能力模型。MCQS将能力建模为结果上的条件概率分布,并将能力学习视为对策略的主动学习问题。该方法利用蒙特卡洛树搜索生成查询,诱导出在极端假设模型(即最悲观与最乐观模型)之间具有高区分价值的执行轨迹。通过执行这些查询,获得信息丰富的状态-动作轨迹,从而高效剔除不一致假设。我们在多个黑箱AI系统上验证,证明了在标准可实现性和采样假设下,该方法具备完备性、一致性与收敛性。实验表明,相较于基线查询策略,MCQS能更高效地学习到准确的能力模型。

原文摘要 · Abstract (English)

Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires methods for characterizing what such systems can do, when they can do it, and what outcomes may result. We introduce Monte Carlo Query Synthesis (MCQS), an active query-synthesis method for learning symbolic stochastic capability models of BBAIs. MCQS models capabilities as conditional probability distributions over outcomes and formulates capability learning as an active learning problem over policies. Our approach uses Monte Carlo tree search to synthesize queries that induce BBAI execution trajectories with high discriminative value between extremal hypothesis models: the lattice meet and join corresponding to the most pessimistic and optimistic hypotheses consistent with the observations. Executing these queries with the agent yields information-rich state-action trajectories that speed up learning by pruning inconsistent hypotheses. We prove soundness, completeness, and convergence properties under standard realizability and sampling assumptions. Experiments with multiple BBAI systems show that MCQS learns accurate capability models more efficiently than baseline query strategies.

AI评估主动学习黑箱系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。