arXiv:2602.04243cs.RO2026-02

让机器人像人一样动态调整视角,提升单摄像头操作能力。

Viewpoint Matters: Dynamically Optimizing Viewpoints with Masked Autoencoder for Visual Manipulation

  • 用预训练的多视角掩码自编码器动态选最佳拍摄角度。
  • 在多个任务中表现优于固定视角,部分超越多摄像头系统。
  • 无需标注视角,适合资源有限的单摄像头机器人场景。

机器人操作仍具挑战性,模仿学习(IL)使机器人可通过专家示范习得任务。现有方法通常依赖固定相机布局,相机位置手动设定且静态不变,严重限制了适应性和覆盖范围。受人类主动感知启发——人类会动态调整视角以获取最相关、最清晰的信息,我们提出MAE-Select框架,用于单摄像头机器人的主动视角选择。该方法充分利用预训练的多视角掩码自编码器表征,无需标签即可在每个时间片段动态选择最具信息量的下一个视角。大量实验表明,MAE-Select显著提升了单摄像头系统的性能,在某些任务中甚至超过多摄像头设置。项目地址:https://mae-select.github.io。

原文摘要 · Abstract (English)

Robotic manipulation continues to be a challenge, and imitation learning (IL) enables robots to learn tasks from expert demonstrations. Current IL methods typically rely on fixed camera setups, where cameras are manually positioned in static locations, imposing significant limitations on adaptability and coverage. Inspired by human active perception, where humans dynamically adjust their viewpoint to capture the most relevant and least noisy information, we propose MAE-Select, a novel framework for active viewpoint selection in single-camera robotic systems. MAE-Select fully leverages pre-trained multi-view masked autoencoder representations and dynamically selects the next most informative viewpoint at each time chunk without requiring labeled viewpoints. Extensive experiments demonstrate that MAE-Select improves the capabilities of single-camera systems and, in some cases, even surpasses multi-camera setups. The project will be available at https://mae-select.github.io.

机器人操作主动感知视角优化自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。