解决机器人手臂被遮挡时动作预测不稳的问题
StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement

- 通过时空特征精炼,融合遮挡掩码与运动连续性约束
- 在严重遮挡下动作准确率提升12.1%,任务成功率提高9.7%
- 适合需要鲁棒视觉动作映射的具身智能系统研发
逆动力学模型(IDM)将视觉观测映射为底层动作指令,是具身人工智能中数据标注与策略执行的核心组件。然而,在机器人手臂部分遮挡这一常见故障模式下,其性能显著下降,导致状态恢复困难和控制不稳定。本文提出StableIDM,一种时空协同的框架,通过三个互补模块增强视觉输入特征:(1) 机器人中心的辅助掩码以抑制背景干扰;(2) 方向特征聚合(DFA),基于可见机械臂方向提取各向异性特征;(3) 时间动态精炼(TDR),利用运动连续性平滑并校正预测结果。大量实验验证:在AgiBot基准上,严重遮挡下严格动作准确率提升12.1%;真实机器人回放任务成功率平均提高9.7%;解码视频生成计划时,端到端抓取成功率提升11.5%;作为自动标注器时,下游视觉-语言-动作模型的真实机器人成功率提升17.6%。结果表明,StableIDM为具身智能中的策略执行与数据生成提供了鲁棒且可扩展的骨干支持。
原文摘要 · Abstract (English)
Inverse Dynamics Models (IDMs) map visual observations to low-level action commands, serving as central components for data labeling and policy execution in embodied AI. However, their performance degrades severely under manipulator truncation, a common failure mode that makes state recovery ill-posed and leads to unstable control. We present StableIDM, a spatio-temporal framework that refines features from visual inputs to stabilize action predictions under such partial observability. StableIDM integrates three complementary components: (1) auxiliary robot-centric masking to suppress background clutter, (2) Directional Feature Aggregation (DFA) for geometry-aware spatial reasoning, which extracts anisotropic features along directions inferred from the visible arm and (3) Temporal Dynamics Refinement (TDR) to smooth and correct predictions via motion continuity. Extensive evaluations validate our approach: StableIDM improves strict action accuracy by 12.1% under severe truncation on the AgiBot benchmark, and increases average task success by 9.7% in real-robot replay. Moreover, it boosts end-to-end grasp success by 11.5% when decoding video-generated plans, and improves downstream VLA real-robot success by 17.6% when functioning as an automatic annotator. These results demonstrate that StableIDM provides a robust and scalable backbone for both policy execution and data generation in embodied artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。