arXiv:2605.07675cs.AIcs.LG2026-05被引 1

构建工业机器人理解评估基准,测试模型因果推理能力。

FactoryBench: Evaluating Industrial Machine Understanding

论文配图:FactoryBench: Evaluating Industrial Machine Understanding
图 1 · 摘自论文原文
  • 基于因果阶梯设计四层问答框架,覆盖状态到决策
  • 超7万条问答数据,零样本下大模型决策准确率不足18%
  • 适合研究工业AI、因果推理与多模态理解的学者

我们提出FactoryBench,一个用于评估时间序列模型和大语言模型在工业机器人遥测数据上机器理解能力的基准。问答对按照佩尔的因果阶梯分为四个层次(状态、干预、反事实、决策),涵盖五种回答格式:四种结构化格式采用确定性评分,自由文本则通过大模型作为裁判的投票机制评分。我们提出一种基于结构化模板的可扩展问答生成框架,发布FactoryWave数据集(来自UR3协作机器人和KUKA KR10工业臂的密集多任务多变量传感器数据),并基于FactoryWave、AURSAD和voraus-AD约1.5万个归一化片段构建了包含超过7万条问答的大规模基准。对六种前沿大模型的零样本评估显示,无模型在结构化层级得分超过50%,决策层面不足18%,揭示当前模型与实际机器理解能力之间存在显著差距。

原文摘要 · Abstract (English)

We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding.

机器理解因果推理工业AI大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。