arXiv:2607.09789cs.AIcond-mat.mtrl-sci2026-07

用自然语言生成辐射输运模拟,靠执行成功率评估效果。

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

论文配图:PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
图 1 · 摘自论文原文
  • 构建包含282个任务的基准,按编辑、修复、复现三类流程评估模型表现。
  • 无领域知识时复现任务成功率为0%,加入知识库后提升至57%。
  • 未来突破需依赖可机器读取的知识库与执行验证环境,不只靠大模型本身。

我们提出PHITSBench,一个针对蒙特卡洛粒子与重离子输运系统(PHITS)的执行评分基准。该基准包含282个可计算输运结果的任务,覆盖参数编辑(Edit)、语法修复(Repair)和从自然语言描述生成完整模拟(Reproduce)三类常见工作流。每个任务通过综合指标得分进行评估,结合执行成功率与生成结果与参考值在输运可观测量上的吻合度。我们使用PHITSBench评估了五种基于GPT-5.4的配置,从零样本提示到知识增强和代理式工作流。未引入领域知识时,模型在编辑与修复任务上分别达到95%和70%的成功率,但在从零开始生成模拟的复现任务中成功率为0%。引入结构化、机器可读的PHITS知识目录后,单次提示的复现任务成功率提升至57%。代理式执行进一步将成功率提升至66%-73%,但带来更高计算开销。失败分析表明,剩余错误主要源于物理可观测量的选择与配置不当,而非语法生成错误。结果表明,未来AI辅助辐射输运建模的进步,将不仅依赖基础模型发展,更取决于可机器读取的知识库、领域训练数据集以及以执行为基准的评估环境。

原文摘要 · Abstract (English)

We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench comprises 282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit), syntax repair (Repair ), and complete simulation generation from natural-language descriptions (Reproduce). Each task is evaluated using a Composite Metric Score that combines execution success with agreement between generated and reference transport observables. Using PHITSBench, we evaluate five GPT-5.4-based configurations ranging from zero-shot prompting to knowledge-augmented and agentic workflows. Without domain-specific knowledge, the model performs well on editing and repair tasks (95% and 70% success, respectively) but fails to generate correct simulations from scratch (0% success on the Reproduce track). A structured, machine-readable PHITS knowledge catalog, supplied alongside the user manual, raises single-shot Reproduce-task success to 57%. Agentic execution provides a further improvement to 66-73%, but at increased computational cost. Failure analysis shows that the remaining errors are dominated by incorrect selection and configuration of physical observables rather than syntax generation. These results suggest that future progress in AI-assisted radiation-transport modeling will depend as much on machine-readable knowledge bases, curated domain-training datasets, and execution-grounded evaluation environments as on advances in foundation models themselves.

辐射输运自然语言生成基准测试AI辅助建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。