arXiv:2508.05186cs.ROcs.CV2025-08中稿 · CVPR被引 5

让机器人自主选视角,提升复杂任务下的操作成功率。

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

  • 学习动态选择对任务有用的最佳虚拟视角。
  • 在真实机器人上实现90%以上成功率,优于现有方法。
  • 适合需要跨任务适应和抗干扰的智能机械臂场景。

当前多任务机器人操作的视觉-语言-动作(VLA)模型通常依赖固定摄像头和共享视觉编码器,导致在遮挡或跨任务迁移时表现受限。为此,我们提出任务感知虚拟视角探索(TVVE)框架,通过学习选择任务相关的虚拟摄像机视角,并利用重建的场景表示动态重渲染观测。为高效实现视角选择,我们在伪环境中训练探索策略;同时引入任务感知专家混合(TaskMoE)视觉编码器,将特征路由至任务专用专家,减少多任务学习中的干扰。为评估分布外鲁棒性,我们构建了带有视觉扰动和相机姿态变化的RLBench-OG基准。在RLBench和RLBench-OG上的实验表明,TVVE显著优于强基线,真实机器人实验进一步验证其对视觉干扰和未见指令的鲁棒性。代码与可视化见:https://hcplab-sysu.github.io/TAVP。

原文摘要 · Abstract (English)

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE), a framework that learns to select task-relevant virtual camera viewpoints and dynamically re-render observations from a reconstructed scene representation using the selected viewpoints. To enable efficient view selection, we train an exploration policy in a pseudo-environment. In addition, we introduce a Task-aware Mixture-of-Experts (TaskMoE) visual encoder that routes visual features to task-specialized experts, mitigating interference in multi-task learning. To evaluate robustness under distribution shifts, we construct RLBench-OG, an out-of-distribution benchmark with visual perturbations and camera pose variations. Experiments on RLBench and RLBench-OG demonstrate that TVVE achieves higher success rates than strong baselines, while real-robot experiments further confirm its robustness to visual disturbances and unseen instructions. Code and visualizations are available at: https://hcplab-sysu.github.io/TAVP.

机器人操作视觉-语言动态视角多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。