用兴趣区域技术让机器人视觉数据跨设备通用,降低采集成本。
ROI-Driven Foveated Attention for Unified Egocentric Representations in Vision-Language-Action Systems
- 通过运动学投影生成手部为中心的兴趣区域,无需额外摄像头。
- 保留关键接触区域的高密度信息,同时保持全局上下文。
- 适合需要多机器人数据复用的具身智能研究者。
具身人工智能系统的发展日益受限于物理交互数据的可获得性与结构。尽管视觉-语言-动作(VLA)模型取得进展,现有流程仍存在数据采集成本高、跨具身对齐有限、互联网规模视觉数据难以迁移到机器人控制等问题。本文提出一种基于兴趣区域(ROI)驱动的工程化工作流,引入以第一人称视角、几何为基准的数据表示。通过正向运动学(FK)将末端执行器姿态投影至单个外部相机,生成与动作对齐的手部中心ROI,无需腕部摄像头或多视角系统。不同于直接下采样全图,本方法在缩放前从原始图像中裁剪出ROI,保留接触关键区域的高局部信息密度,同时维持全局上下文。我们构建了可复现的全流程,涵盖标定、同步、ROI生成、确定性边界处理和元数据治理。所得表示具有具身对齐与视角归一化特性,支持异构机器人间的数据复用。我们认为,第一人称视角的ROI是一种实用的数据抽象,有助于实现大规模数据采集与跨具身学习,弥合互联网规模感知与机器人特异性控制之间的鸿沟。
原文摘要 · Abstract (English)
The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data collection cost, limited cross-embodiment alignment, and poor transfer from internet-scale visual data to robot control. We propose a region-of-interest (ROI) driven engineering workflow that introduces an egocentric, geometry-grounded data representation. By projecting end-effector poses via forward kinematics (FK) into a single external camera, we derive movement-aligned hand-centric ROIs without requiring wrist-mounted cameras or multi-view systems. Unlike directly downsampling the full frame, ROI is cropped from the original image before resizing, preserving high local information density for contact-critical regions while retaining global context. We present a reproducible pipeline covering calibration, synchronization, ROI generation, deterministic boundary handling, and metadata governance. The resulting representation is embodiment-aligned and viewpoint-normalized, enabling data reuse across heterogeneous robots. We argue that egocentric ROI serves as a practical data abstraction for scalable collection and cross-embodiment learning, bridging internet-scale perception and robot-specific control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。