arXiv:2606.10359cs.AI2026-06

让大模型理解供应链政策,同时保持物理合理性。

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience

论文配图:ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience
图 1 · 摘自论文原文
  • 构建融合物理规律的供应链世界模型,用6维隐空间表示网络。
  • 在半导体供应链测试中,推理一致性提升33%,抗冲击能力达82.3%。
  • 适合研究智能供应链、多模态决策系统的学者和工程师。

供应链中的智能体面临根本性的认知鸿沟:大语言模型(LLMs)能解读政策但缺乏物理依据,强化学习(RL)可优化物流流但无法理解非结构化约束。我们提出REFLECTICHAIN,通过生成式供应链世界模型(SC-WM)将异构供应链网络编码为包含物理守恒关系的6维图-隐空间,并采用双环学习机制,分离认知不确定性(基于KL信任域的策略调整)与随机不确定性(随机隐状态回溯)。在包含10个节点、SIR风险传播机制、6种扰动类型及10种策略约束模板的半仿真基准(Semi-Sim)上,REFLECTICHAIN使推理一致性得分提升33.0%(p < 0.0001, d = 2.78),在对抗性冲击下维持82.3%的可操作性,并表现出反脆弱性(中等压力下收益增加40.2%)。我们识别出三种操作认知机制——不确定性分离、知识边界检测与经验贝叶斯策略更新,并讨论五类局限性。

原文摘要 · Abstract (English)

AI agents in supply chains face a fundamental epistemic gap: large language models (LLMs) interpret policies but lack physical grounding, while reinforcement learning (RL) optimizes flows but is semantically blind to unstructured constraints. We introduce REFLECTICHAIN, bridging this gap through a Generative Supply Chain World Model (SC-WM) - encoding heterogeneous supply networks into a 6-dim graph-latent space with physical conservation - and Double-Loop Learning that separates epistemic uncertainty (KL-trust-region-bounded policy adaptation) from aleatoric uncertainty (stochastic latent rollouts). On Semi-Sim, a 10-node semiconductor benchmark with SIR risk propagation, 6 perturbation types, and 10 policy constraint templates, REFLECTICHAIN improves Rationale Consistency Score by 33.0% (p < 0.0001, d = 2.78), maintains 82.3% operability under adversarial shocks, and exhibits anti-fragile behavior (+40.2% gain under moderate pressure). We identify three operational epistemic mechanisms - uncertainty separation, knowledge-boundary detection, and empirical Bayesian policy updating - and discuss five limitation categories.

供应链大模型世界模型双环学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。