arXiv:2605.28994cs.AI2026-05

构建AI建模评估基准,推动可解释、以人为本的智能建模工具发展。

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

  • 建立开放评测框架,自动化测试AI在建模中的表现
  • 实测显示AI在定性讨论中表现优于因果推理与定量纠错
  • 强调人机协作,适合关注可信AI建模的研究者与开发者

用于支持现实决策的AI工具需能构建可解释的仿真模型。为确保工具辅助而非替代人类专业判断,BEAMS倡议旨在通过基准测试引导负责任且合乎伦理的建模与仿真工具发展。该倡议利用开源数字与组织基础设施,协同评估建模与仿真类AI工具。由倡议托管的开源sd ai项目保障透明度,并促进成果共享。指导小组聚焦优先级设定,技术小组则将基准转化为自动化测试。目前已实现对定性建模、定量建模及模型讨论等多类任务的测评,涵盖因果转化、模型迭代、因果推理、一致性、行为解释、建模建议步骤与修复建议等维度。当sd ai引擎与不同LLM结合时,性能差异明显。测评表明,当前AI建模工具在讨论与基础定性任务上表现较好,但在因果推理与定量错误修复方面仍不足。无单一LLM在所有引擎类型中占优,凸显任务特异性及速度与准确率间的权衡。倡议持续拓展基准,以纳入偏见考量与以人为本的使用场景。

原文摘要 · Abstract (English)

AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable. Tools that can automate aspects of modeling practice must complement human expertise, not replace it. The BEAMS Initiative aims to guide the development of AI tools for modeling and simulation toward forms that are responsible and ethical by establishing benchmarks for human centered modeling and simulation practices. The initiative uses open digital and organizational infrastructure to collaboratively evaluate AI tools for modeling and simulation. The open source sd ai project hosted by the initiative establishes transparency and enables contributions to be shared broadly. A steering group focuses on prioritizing potential benchmarks, while a technical group focuses on implementing the benchmarks in the form of automated tests. Tests for several distinct categories of evaluation have been implemented and applied to AI tools that support qualitative model building, quantitative model building, and model discussion. These include tests for causal translation, model iteration, causal reasoning, conformance, model behavior explanation, suggested model building steps, and suggested model fixes. When engines from the sd ai project are coupled with different LLMs, their performance on these evaluations reveals variability across different AI tools. The evaluations implemented by the initiative demonstrate that AI enabled modeling tools perform better at discussion and basic qualitative tasks than with causal reasoning and quantitative error fixing. No single LLM dominates across engine types, highlighting the importance of specific tasks and tradeoffs between speed and accuracy. Ongoing efforts of the initiative aim to incorporate benchmarks that address concerns about bias by considering alternative perspectives and human centered use cases.

AI建模评估基准人机协作可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。