arXiv:2608.21402cs.ROcs.CV2026-08

提出只对视角不变部分施加一致性损失,提升机器人动作模型在未知视角下的鲁棒性。

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

论文配图:Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information
图 1 · 摘自论文原文
  • 仅对动作、本体感知等视角不变量施加跨视角一致性约束
  • 在未见轨道视角下闭环成功率提升12.2点,置信区间[7.4,17.0]
  • 无需相机位姿、深度或视图合成,部署接口不变,适合真实场景应用

世界动作模型(WAMs)联合去噪未来视频帧与机器人动作,视频先验期望其控制具有泛化能力。摄像机视角变化仍是其最严峻的扰动因素。本文研究该模型类的特定问题:当使用同状态跨视角图像对训练时,一致性损失应作用于输出的哪些坐标?由于去噪目标混合了视角相关(如未来场景)与视角无关(如动作块、未来本体感知、价值)的坐标,我们证明对相关部分施加一致性损失会严重削弱合法视角特异性内容,使其压缩至真实值的$1/(1+4λ)$。因此,选择性跨视角一致性(SCVC)仅约束不变块,不依赖相机标签、外参、深度或视图合成,训练与测试均无需额外信息,且保持部署接口不变。我们在LIBERO-Plus相机轨道任务上引入切割与留出评估协议,分离分布匹配的上限与真正的内插/外推效果,匹配对训练控制隔离一致性项的影响。在超出训练范围的未见轨道视角下,SCVC相比匹配对照组提升闭环成功率12.2点(95% CI [7.4, 17.0];第二种子独立验证+15.5, CI [11.7, 19.4]),此效果在另两个相机轴上复现;而在训练范围内内插无增益(-1.2、-4.3点),分布内性能保持稳定(-0.6、-0.2)。此外,跨主干审计发现,已有相机鲁棒性指标受腕部相机姿态稳定性干扰。

原文摘要 · Abstract (English)

World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4λ)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.

动作建模视角不变性机器人控制跨视角一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。