通过显式3D参考关系提升扩散策略在视角与位置变化下的泛化能力
MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference

- 将参考轨迹的匹配对转化为显式3D空间关系作为条件输入
- 在联合视角偏移±15°和门位置超出训练范围时仍达96.67%成功率
- 适合需要高鲁棒性视觉动作规划的机器人任务场景
行为克隆的视觉运动策略在训练分布附近表现准确,但在物体位置与相机视角同时变化时会失效。一个有效的参考轨迹包含可迁移交互所需的几何信息,但策略需将该几何与当前场景对齐,并在去噪过程中保持敏感性。为此,我们提出 MemCorr-DP,一种将冻结的 RoMa v2 匹配对提升为当前场景与参考轨迹间显式3D关系的扩散策略。反事实成对目标在相同物理状态和噪声动作下赋予相反行为,同时保留参考特定的去噪目标。混合条件微调使策略从真实几何过渡到测量到的对应误差。最强评估设置中,门位于训练支持范围外的最外侧区域,查询相机视角改变±15°。在此联合偏移下,MemCorr-DP 达到96.67%闭环成功率,优于使用相同动作架构的视觉Transformer的88.00%。消融实验与参考干预表明行为响应所选参考,而对照组显示完整关系集优于仅未来运动或质心几何。结果支持显式3D参考关系在空间与视角复合变化任务中作为稳健条件接口。
原文摘要 · Abstract (English)
Behavior-cloned visuomotor policies can remain accurate near their training distribution yet fail when object position and camera viewpoint change together. A successful reference trajectory contains the geometry needed to transfer the same interaction, but the policy must align that geometry with the current scene and remain sensitive to it during denoising. To address these challenges, we present MemCorr-DP, a diffusion policy that lifts frozen RoMa v2 matches into explicit 3D relations between the current scene and the reference trajectory. A counterfactual paired objective assigns opposite behaviors the same physical state and noisy action while retaining reference-specific denoising targets. Mixed-condition fine-tuning then adapts the policy from ground-truth geometry to measured correspondence errors. Our strongest evaluation places the Door in the outermost position bands beyond the training support and changes the query camera by $\pm15^\circ$. Under this combined shift, MemCorr-DP achieves 96.67% closed-loop success, compared with 88.00% for a visual Transformer with the same action architecture. Objective ablations and reference interventions show that behavior responds to the selected reference, while matched controls favor the complete relation set over future motion or centroid geometry alone. These results support explicit 3D reference relations as a robust conditioning interface when spatial and viewpoint changes are compounded in the evaluated task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。