用视觉语言动作模型预判机器人动作风险与任务进展,实现安全高效部署。
Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields
- 基于动作条件的JEPA模型在冻结编码空间中预测动作块的进展与风险
- 在600次仿真中提升成功率,同时降低碰撞漏检率至相同召回率下
- 适合需要高安全性且适配多种机器人的自主系统研发者
视觉-语言-动作策略泛化能力强但无执行时保障;经典模型规划方法满足运动与几何约束但泛化差。本文研究是否可通过动作条件的联合嵌入预测架构(JEPA)世界模型,在动作执行前预判候选动作块的任务进展与物理风险,并结合针对具体机器人形态的模型基安全屏障,构建适用于异构机器人的可部署流程。提出一种滚动时域决策框架:(1) 候选生成器产生K个候选动作块;(2) 动作条件的JEPA在冻结编码器隐空间中对每个候选进行前向推演,受机器人形态嵌入条件控制;(3) 校准的风险与进展头评分每条推演结果并报告置信度;(4) 针对每种机器人形态的确定性安全屏障过滤不可行候选;(5) 备用阶梯机制处理无可行候选情形。学习到的排序仅重排可行候选,执行保障由确定性屏障与备用阶梯提供。在模拟环境(LIBERO-Long)中按预注册协议评估,600期配置下,全框架相较仅使用屏障的基线提升成功率,同时在匹配召回率下减少碰撞漏报。包含目标机器人及边缘加速器上的部署效率测量。真实机器人实验与离线重排序显著性检验为后续工作;详见论文披露。
原文摘要 · Abstract (English)
Vision-language-action policies generalize broadly but provide no execution-time guarantees; classical model-based planners respect kinematic and geometric constraints but generalize poorly. We study whether an action-conditioned Joint-Embedding Predictive Architecture (JEPA) world model can predict, before execution, both task progress and physical risk for candidate action chunks, and whether coupling these predictions to an embodiment-specific model-based safety shield yields a deployable pipeline for heterogeneous robots. We propose a receding-horizon decision pipeline: (1) a proposer produces K candidate action chunks; (2) an action-conditioned JEPA rolls each candidate forward in a frozen-encoder latent space conditioned on an embodiment embedding; (3) calibrated risk and progress heads score each rollout and report uncertainty; (4) a deterministic per-embodiment safety shield filters inadmissible candidates; (5) a fallback ladder handles empty-admissible-set cases. The learned ranking only reorders admissible candidates; enforcement guarantees come from the deterministic shield and fallback ladder. We evaluate with a pre-registered protocol in simulation (LIBERO-Long). In 600-episode configurations the full framework improved success over a shield-only baseline and reduced collision false negatives at matched recall. Deployment-efficiency measurements on target on-robot and edge accelerators are included. Real-robot experiments and an offline reranking significance test remain future work; see the paper for disclosures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。