通过单帧编辑与机器人状态,高效增强4D多视角机器人图像数据。
ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents
- 基于单帧编辑和机器人状态,统一处理多视角时序数据。
- 在模拟与真实环境上显著提升VLA模型的泛化能力与鲁棒性。
- 适合需要高质量视觉-语言-动作训练数据的研究者使用。
机器人模仿学习依赖4D多视角序列图像,但数据采集成本高且高质量数据稀缺,严重制约了具身智能策略(如视觉-语言-动作模型)的泛化与应用。数据增强是缓解数据稀缺的有效策略,但针对操作任务的4D多视角序列图像编辑方法仍为空白。为此,我们提出ERMV(Editing Robotic Multi-View 4D data),一种新型数据增强框架,可基于单帧编辑与机器人状态高效编辑整段多视角序列。该任务面临三大挑战:(1) 动态视图与长时序下的几何与外观一致性;(2) 在低计算成本下扩展编辑窗口;(3) 关键物体(如机械臂)的语义完整性。ERMV通过三项创新应对:首先,引入新提出的视点运动感知注意力(EMA-Attn)机制,在施加几何约束前学习运动引起的像素位移,以保证运动模糊下的时空一致性;其次,首创稀疏时空(STT)模块,解耦时空维度,通过稀疏采样重构单帧多视角问题,降低计算开销;第三,引入反馈干预机制,利用多模态大语言模型(MLLM)检测编辑不一致,并仅在必要时请求专家指导,缓解误差累积。大量实验表明,ERMV增强的数据显著提升了VLA模型在模拟与真实环境中的鲁棒性与泛化性能。
原文摘要 · Abstract (English)
Robot imitation learning relies on 4D multi-view sequential images. However, the high cost of data collection and the scarcity of high-quality data severely constrain the generalization and application of embodied intelligence policies like Vision-Language-Action (VLA) models. Data augmentation is a powerful strategy to overcome data scarcity, but methods for editing 4D multi-view sequential images for manipulation tasks are currently lacking. Thus, we propose ERMV (Editing Robotic Multi-View 4D data), a novel data augmentation framework that efficiently edits an entire multi-view sequence based on single-frame editing and robot state conditions. This task presents three core challenges: (1) maintaining geometric and appearance consistency across dynamic views and long time horizons; (2) expanding the working window with low computational costs; and (3) ensuring the semantic integrity of critical objects like the robot arm. ERMV addresses these challenges through a series of innovations. First, to ensure spatio-temporal consistency in motion blur, we introduce a novel Epipolar Motion-Aware Attention (EMA-Attn) mechanism that learns pixel shift caused by movement before applying geometric constraints. Second, to maximize the editing working window, ERMV pioneers a Sparse Spatio-Temporal (STT) module, which decouples the temporal and spatial views and remodels a single-frame multi-view problem through sparse sampling of the views to reduce computational demands. Third, to alleviate error accumulation, we incorporate a feedback intervention Mechanism, which uses a Multimodal Large Language Model (MLLM) to check editing inconsistencies and request targeted expert guidance only when necessary. Extensive experiments demonstrate that ERMV-augmented data significantly boosts the robustness and generalization of VLA models in both simulated and real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。