arXiv:2608.10030cs.AIcs.MA2026-08

用自动化系统AEROBAT研究AI智能体行为,效率远超人工。

Automating and Scaling Behavioral Scientific Research on AI Agents

论文配图:Automating and Scaling Behavioral Scientific Research on AI Agents
图 1 · 摘自论文原文
  • 通过自动生成假设、设计实验、分析结果完成完整研究流程
  • 测试73个假说,执行2.3万次模拟,发现30个有统计意义的结果
  • 适合想高效探索AI行为机制的研究者和开发者

随着AI智能体在复杂环境中的广泛应用,理解其行为变得至关重要。然而,对AI智能体的行为科学研究仍以人工为主,耗时费力。我们提出AEROBAT,首个可自动化进行行为科学研究的多智能体系统。用户指定目标行为后,AEROBAT能自动完成从假设生成、受控实验设计与执行、行为评估、结果分析到报告撰写的完整研究流程。针对12种目标行为,共生成并测试了73个假设,设计了1,160个受控实验,累计执行22,954次模拟。其中30个假设获得中等至强统计证据支持,部分为新发现。结果表明,自动化行为科学研究可有效补充并扩展人工研究的范围与深度。

原文摘要 · Abstract (English)

As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence was found for 30 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.

AI行为自动化研究多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。