arXiv:2604.04967cs.ROcs.LG2026-04被引 1

提出轻量级信念追踪模块,实时检测协作机器人行为突变,显著降低碰撞风险。

Belief Dynamics for Detecting Behavioral Shifts in Safe Collaborative Manipulation

  • 用选择性状态空间动态与因果注意力构建信念追踪机制。
  • 在5种子实验中实现85.7%的检测率(容忍±3步),关闭距离仅4.8步。
  • 不修改主策略,仅增加7.4毫秒开销,适合部署于现有机器人系统。

在共享工作空间中,协作机器人需应对其他智能体行为随时间变化的问题。当合作方中途改变策略时,沿用旧假设可能导致危险动作和更高碰撞风险。本文在ManiSkill任务中研究受控非平稳性下的策略切换检测,对比十种方法与五组随机种子,启用检测可使切换后碰撞减少52%。然而平均表现掩盖了可靠性差异:在±3步容忍度下,检测率从86%降至30%,而±5步时所有方法均达100%。提出UA-TOM,一种轻量级信念追踪模块,通过选择性状态空间动力学、因果注意力和预测误差信号增强冻结的视觉-语言-动作(VLA)控制骨干。在五种子与1200个回合中,其在无辅助方法中表现最佳(±3步下检测率达85.7%),关闭距离仅4.8步,优于基准(5.3步)。分析显示,状态更新幅度在切换时提升17倍,衰减周期约10个时间步,离散化步长收敛至Δ_t ≈ 0.78,表明敏感性源于学习到的动力学而非输入门控。跨域实验(Overcooked)验证因果注意力与预测误差信号的互补作用。UA-TOM仅引入7.4毫秒推理延迟(占50毫秒控制预算的14.8%),可在不修改基线策略前提下实现可靠检测。

原文摘要 · Abstract (English)

Robots operating in shared workspaces must maintain safe coordination with other agents whose behavior may change during task execution. When a collaborating agent switches strategy mid-episode, continuing under outdated assumptions can lead to unsafe actions and increased collision risk. Reliable detection of such behavioral regime changes is therefore critical. We study regime-switch detection under controlled non-stationarity in ManiSkill shared-workspace manipulation tasks. Across ten detection methods and five random seeds, enabling detection reduces post-switch collisions by 52%. However, average performance hides significant reliability differences: under a realistic tolerance of +-3 steps, detection ranges from 86% to 30%, while under +-5 steps all methods achieve 100%. We introduce UA-TOM, a lightweight belief-tracking module that augments frozen vision-language-action (VLA) control backbones using selective state-space dynamics, causal attention, and prediction-error signals. Across five seeds and 1200 episodes, UA-TOM achieves the highest detection rate among unassisted methods (85.7% at +-3) and the lowest close-range time (4.8 steps), outperforming an Oracle (5.3 steps). Analysis shows hidden-state update magnitude increases by 17x at regime switches and decays over roughly 10 timesteps, while the discretization step converges to a near-constant value (Delta_t approx 0.78), indicating sensitivity driven by learned dynamics rather than input-dependent gating. Cross-domain experiments in Overcooked show complementary roles of causal attention and prediction-error signals. UA-TOM introduces 7.4 ms inference overhead (14.8% of a 50 ms control budget), enabling reliable regime-switch detection without modifying the base policy.

机器人协作行为检测信念追踪安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。