用事件流提升视觉语言动作模型在光照变化下的鲁棒性
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

- 通过动作查询路由机制融合事件流信息,增强光照变化下的感知
- 在低光和近黑暗环境中成功率显著提升,仿真与真实世界均验证有效
- 适合需要跨光照环境稳定执行任务的机器人应用
视觉-语言-动作(VLA)模型已成为具身智能的重要范式。然而现有模型通常假设光照良好且稳定的室内环境,而真实世界的具身操作常面临光照变化导致的视觉退化问题,严重威胁机器人操作的鲁棒性。为此,我们提出Event-VLA,一种面向不同光照条件下的通用操作任务的事件增强型VLA框架。将光照退化下的VLA操作建模为以RGB为中心策略的实用性鲁棒性问题,引入事件流作为对光照不敏感、对运动敏感的互补观测信号,以提升不同可见度下的鲁棒性。不同于传统多模态融合直接将事件特征注入全局语义令牌空间,Event-VLA通过动作查询路由路径注入事件信息:利用可学习的动作查询从VLA推理过程中提取任务相关语义,并通过门控交叉注意力选择性聚合事件令牌,构建事件感知的动作表示。该设计保留预训练的RGB-语言语义先验,同时有效利用事件信息实现鲁棒动作预测。仿真与真实世界部署实验表明,Event-VLA在正常光照下保持强操作性能,并在低光退化及近黑暗真实场景中显著提升成功率。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-world embodied manipulation may involve degraded RGB observations caused by illumination shifts, posing critical challenges for robust robotic manipulation. To address this gap, we propose \textbf{Event-VLA}, an event-enhanced VLA framework for generalizable manipulation across varying illumination conditions. We formulate VLA-based manipulation under degraded visibility as a practical robustness problem for RGB-centric policies, and introduce event streams as an illumination-robust, motion-sensitive complementary observation to improve robustness across visibility levels. Specifically, unlike conventional multimodal fusion that directly merges event features into the global semantic token space, Event-VLA injects event information through an action-query routing pathway. It uses learnable action queries to extract task-relevant semantics from the VLA reasoning process, and selectively aggregates event tokens via gated cross-attention to construct event-aware action representations. This design preserves the pretrained RGB-language semantic priors while effectively leveraging event information for robust action prediction. Experiments in simulation and real-world deployment show that Event-VLA maintains strong manipulation performance under normal lighting and improves success rates under low-light degradation and near-dark real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。