arXiv:2603.09053cs.LGcs.AI2026-03被引 1

提升仿真到决策的鲁棒性,让模型在不安全区域也能做出稳定可靠决策。

Sim2Act: Robust Simulation-to-Decision Learning via Adversarial Calibration and Group-Relative Perturbation

  • 通过对抗校准重加权关键状态动作对的误差,使仿真精度与决策影响对齐。
  • 采用组相对扰动策略,在不确定环境下稳定策略学习,避免过度保守。
  • 适用于供应链等高风险场景,适合需要稳定决策的工业系统部署。

仿真到决策学习可在数字环境中安全训练策略,避免真实世界部署风险,已在供应链和工业系统等关键领域广泛应用。然而,基于噪声或偏差真实数据训练的仿真器常在决策敏感区域出现预测误差,导致动作排序不稳定和策略不可靠。现有方法要么只关注平均仿真保真度,要么采用保守正则化,可能因剔除高风险高回报动作而引发策略坍缩。本文提出Sim2Act框架,同时提升仿真器与策略的鲁棒性:首先引入对抗校准机制,对决策关键的状态-动作对进行误差重加权,使代理保真度与下游决策影响对齐;其次设计组相对扰动策略,在仿真不确定性下稳定策略学习,不施加过度悲观约束。在多个供应链基准测试中,该方法在结构化与非结构化扰动下均表现出更优的仿真鲁棒性和更稳定的决策性能。

原文摘要 · Abstract (English)

Simulation-to-decision learning enables safe policy training in digital environments without risking real-world deployment, and has become essential in mission-critical domains such as supply chains and industrial systems. However, simulators learned from noisy or biased real-world data often exhibit prediction errors in decision-critical regions, leading to unstable action ranking and unreliable policies. Existing approaches either focus on improving average simulation fidelity or adopt conservative regularization, which may cause policy collapse by discarding high-risk high-reward actions. We propose Sim2Act, a robust simulation-to-decision framework that addresses both simulator and policy robustness. First, we introduce an adversarial calibration mechanism that re-weights simulation errors in decision-critical state-action pairs to align surrogate fidelity with downstream decision impact. Second, we develop a group-relative perturbation strategy that stabilizes policy learning under simulator uncertainty without enforcing overly pessimistic constraints. Extensive experiments on multiple supply chain benchmarks demonstrate improved simulation robustness and more stable decision performance under structured and unstructured perturbations.

仿真学习供应链鲁棒决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。