arXiv:2601.17885cs.CVcs.AI2026-01中稿 · ICRA被引 6

让双臂机器人在复杂场景中更稳地操作,通过增强视觉语言动作模型的感知与理解。

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

  • 分 token 预测深度分布,融合多视角局部信息形成几何对齐表征
  • 在随机化环境中成功率达 82.7%,比最强基线提升 23.0 个百分点
  • 适合需要鲁棒感知与精细指令理解的双臂机器人任务研究者

在杂乱场景中进行双臂操作需策略在遮挡、视角变化和场景差异下保持稳定。现有视觉-语言-动作模型常因(i)多视角特征通过无视图感知的令牌拼接融合,导致跨视角空间表征有限;(ii)语言作为全局条件注入,造成指令定位粗糙。本文提出 PEAfowl,一种感知增强的多视角 VLA 策略用于双臂操作。在空间感知方面,PEAfowl 预测每个令牌的深度分布,执行可微 3D 提升,并聚合本地跨视角邻居,形成几何对齐的跨视角对齐表征。在语言利用方面,采用类似 Perceiver 的文本感知读出机制,在冻结的 CLIP 视觉特征上迭代积累证据,替代全局条件。为更好利用商品级 RGB-D 传感器(尽管深度数据噪声大、不完整),其深度分布提升支持仅训练阶段的深度蒸馏,由预训练深度教师监督深度分布头,注入精炼几何先验而无需增加推理开销。在域随机化的 RoboTwin 2.0 设置下,PEAfowl 将最强基线的成功率提升 23.0 个百分点至 82.7%,物理实验也验证了其在真实机器人任务中的性能优势。

原文摘要 · Abstract (English)

Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint changes and scene variations. Existing vision-language-action models often lack such robustness because (i) multi-view features are fused via view-agnostic token concatenation, yielding limited cross-view spatial representations, and (ii) language is injected as global conditioning, resulting in coarse instruction grounding. In this paper, we introduce PEAfowl, a perception-enhanced multi-view VLA policy for bimanual manipulation. For spatial perception, PEAfowl predicts per-token depth distributions, performs differentiable 3D lifting, and aggregates local cross-view neighbors to form geometrically grounded, cross-view aligned representations. For language utilization, we propose to replace global conditioning with a Perceiver-style text-aware readout over frozen CLIP visual features, enabling iterative evidence accumulation. To better exploit commodity RGB-D sensing despite noisy and incomplete depth, PEAfowl's depth-distribution lifting naturally supports training-only depth distillation, where a pretrained depth teacher supervises the depth-distribution head to inject refined geometric priors without adding inference overhead. On RoboTwin 2.0 under domain-randomized setting, PEAfowl improves the strongest baseline by 23.0 pp in success rate, and physical experiments further demonstrate improved performance on the evaluated real-robot tasks. Project website: https://peafowlvla.github.io/.

双臂操作多视角感知视觉语言动作深度蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。