arXiv:2505.18899cs.CVcs.LG2025-05被引 1

用事件驱动视觉替代传统图像,让机器人模仿更抗光照变化。

Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos

  • 将视频转为稀疏事件流,只保留动态变化信息
  • 在多个环境上实现无须数据增强的稳定模仿性能
  • 适合需要跨域鲁棒性的机器人控制场景

从视频中学习模仿在专家演示与学习者环境存在领域差异(如光照、颜色、纹理不同)时往往失效。虽然视觉随机化通过数据增强部分缓解此问题,但计算成本高且对未见场景应对能力弱。本文提出新思路:不随机化外观,而是彻底消除其影响,重新设计感知表示。受生物视觉系统(如视网膜神经节细胞)和新型传感器启发,引入事件驱动感知,将标准RGB视频转化为仅编码时间强度梯度的稀疏事件流,丢弃静态外观特征。该生物启发方法将运动动态与视觉风格解耦,使模仿在专家与代理环境存在视觉不匹配时仍具鲁棒性。在DeepMind Control Suite和Adroit平台的动态灵巧操作任务上验证了有效性。代码已公开于Eb-LAIfO。

原文摘要 · Abstract (English)

Imitation from videos often fails when expert demonstrations and learner environments exhibit domain shifts, such as discrepancies in lighting, color, or texture. While visual randomization partially addresses this problem by augmenting training data, it remains computationally intensive and inherently reactive, struggling with unseen scenarios. We propose a different approach: instead of randomizing appearances, we eliminate their influence entirely by rethinking the sensory representation itself. Inspired by biological vision systems that prioritize temporal transients (e.g., retinal ganglion cells) and by recent sensor advancements, we introduce event-inspired perception for visually robust imitation. Our method converts standard RGB videos into a sparse, event-based representation that encodes temporal intensity gradients, discarding static appearance features. This biologically grounded approach disentangles motion dynamics from visual style, enabling robust visual imitation from observations even in the presence of visual mismatches between expert and agent environments. By training policies on event streams, we achieve invariance to appearance-based distractors without requiring computationally expensive and environment-specific data augmentation techniques. Experiments across the DeepMind Control Suite and the Adroit platform for dynamic dexterous manipulation show the efficacy of our method. Our code is publicly available at Eb-LAIfO.

视觉模仿事件相机领域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。