让视觉语言模型在制造中可靠执行,靠的是能记住状态、会自我诊断的外部世界模型。
VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing
- 用可查询的外部世界模型分离状态记忆与推理,避免状态漂移。
- 故障时通过状态比对精准定位问题,恢复成功率从不足5%提升至95%。
- 适合需要长期稳定运行的智能工厂机器人系统,尤其看重可靠性场景。
视觉语言模型(VLM)在智能制造中的高层规划展现潜力,但在动态工位部署面临两大挑战:(1) 无状态操作,无法持续追踪视野外状态,导致世界状态漂移;(2) 推理不透明,故障难以诊断,引发高成本盲重试。本文提出VLM-DEWM,一种认知架构,通过持久可查询的动态外部世界模型(DEWM)将VLM推理与世界状态管理解耦。每个VLM决策被结构化为可外化的推理轨迹(ERT),包含动作建议、世界信念和因果假设,执行前经由DEWM验证。故障发生时,通过预测与观测状态间的差异分析,实现定向恢复而非全局重规划。我们在多工位装配、大规模设施探索及真实机器人故障恢复任务中评估VLM-DEWM。相比基线记忆增强型VLM系统,其状态追踪准确率从56%提升至93%,恢复成功率从低于5%升至95%,并通过结构化内存显著降低计算开销。结果表明,VLM-DEWM是动态制造环境中长周期机器人操作的可验证、高鲁棒性解决方案。
原文摘要 · Abstract (English)
Vision-language model (VLM) shows promise for high-level planning in smart manufacturing, yet their deployment in dynamic workcells faces two critical challenges: (1) stateless operation, they cannot persistently track out-of-view states, causing world-state drift; and (2) opaque reasoning, failures are difficult to diagnose, leading to costly blind retries. This paper presents VLM-DEWM, a cognitive architecture that decouples VLM reasoning from world-state management through a persistent, queryable Dynamic External World Model (DEWM). Each VLM decision is structured into an Externalizable Reasoning Trace (ERT), comprising action proposal, world belief, and causal assumption, which is validated against DEWM before execution. When failures occur, discrepancy analysis between predicted and observed states enables targeted recovery instead of global replanning. We evaluate VLM-DEWM on multi-station assembly, large-scale facility exploration, and real-robot recovery under induced failures. Compared to baseline memory-augmented VLM systems, VLM DEWM improves state-tracking accuracy from 56% to 93%, increases recovery success rate from below 5% to 95%, and significantly reduces computational overhead through structured memory. These results establish VLM-DEWM as a verifiable and resilient solution for long-horizon robotic operations in dynamic manufacturing environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。