arXiv:2607.22563cs.AI2026-07

用合成场景扩充工业智能体评测基准,提升真实性和扩展性。

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

论文配图:Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
图 1 · 摘自论文原文
  • 构建可验证的合成场景生成管道,融合标准与物理合理性约束
  • 生成50个场景时效率提升8倍,质量评分74.2±1.9,优于基线
  • 适合工业智能体评估、自动化测试与系统验证研究者使用

工业智能体评测需整合遥测数据、故障模式、维护记录和领域标准的真实场景。现有基准如AssetOpsBench依赖人工编写场景,覆盖资产类型有限。本文在AssetOpsBench基础上新增智能电网变压器资产类,并引入四种符合IEC标准的诊断工具:健康指数预测、溶解气体分析、绕组温度评估和负载曲线评估。提出ScenarioGeneratorAgent合成场景生成管道,通过证据驱动的资产画像、覆盖感知的预算分配,以及混合验证-修复循环,确保场景满足结构合法性、工具可达性、物理合理性、标准一致性与去重要求。为提升可扩展性,采用双层缓存、并行焦点组生成、线程池卸载、批量LLM调用与早期拒绝过滤。在智能电网变压器场景生成中,该优化使50个场景的端到端运行时间减少8倍,复合质量得分达74.2±1.9,优于未优化基线的73.8±3.0。结果表明,基于标准的合成场景生成能高效扩展工业智能体基准而不牺牲质量。

原文摘要 · Abstract (English)

Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by $8\times$ for 50 scenarios while preserving quality, achieving a composite quality score of $74.2 \pm 1.9$ compared with $73.8 \pm 3.0$ for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.

工业智能体场景生成合成数据评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。