arXiv:2605.29078cs.AIcs.LG2026-05中稿 · publication at the…

为工业调度中的真实执行与模拟差距提供可追踪的执行层。

Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics

  • 构建事件流驱动的决策快照,统一异步状态
  • 将执行失败分类为可追溯的类型,实现全链路归因
  • 适合需要高可靠性调度系统的工业场景

事件驱动的调度策略在工业环境中日益普及,但系统状态异步且部分可观测,导致决策状态时间不一致、动作合法性未明确定义,执行错误来源模糊,限制了可靠性和可解释性。为此,提出一种与策略无关的执行与度量层,用于连接调度策略与工业执行环境。该层从异步事件流中构建决策有效快照,定义具有明确动作合法性的标准化执行契约,并记录政策意图、事务结果、物理执行与人工干预之间的偏差。这实现了决策语义与执行行为的分离,使部署不匹配可见且可结构化归因。框架在离散事件仿真中评估,结果表明在所有观测延迟条件下均具分析优势:未区分的执行失败被转化为结构化、类型化的输出,具备完整归因覆盖。在低观测延迟下运营收益最强,可提前预防可避免的执行错误。整体上,该层将执行不确定性转化为监督数据,用于评估与策略优化。

原文摘要 · Abstract (English)

Event-driven scheduling policies are increasingly deployed in industrial environments, where decisions are made under asynchronous and partially observed system states. As a result, decision states are not temporally consistent, action admissibility is not explicitly defined, and the origin of execution errors remains ambiguous. These issues limit both reliability and interpretability. To address this gap, a policy-neutral execution and measurement layer is proposed to mediate between scheduling policies and the industrial execution environment. The layer constructs decision-valid snapshots from asynchronous event streams, defines a standardized execution contract with explicit action admissibility, and records outcomes as divergences between policy intent, transactional outcomes, physical execution, and human intervention. This enables a separation between decision semantics and execution behavior and makes deployment mismatch observable and structurally attributable. The proposed framework is evaluated using a discrete-event simulation. The results show analytical benefits across all observation lag regimes, as undifferentiated execution failures are transformed into structured, typed outcomes with full attribution coverage. Operational benefits are strongest under low observation lag, where avoidable execution errors can be prevented before commitment. Overall, the layer turns execution uncertainty into supervisory data for evaluation and policy refinement.

强化学习工业调度执行对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。