arXiv:2409.02669cs.ROcs.AI2024-09被引 8

用因果视角改进机器人导航的Transformer模型,提升泛化能力。

Causality-Aware Transformer Networks for Robotic Navigation

  • 引入因果理解模块,增强环境建模能力
  • 端到端训练,超越多个基准任务表现
  • 适用于强化学习与监督学习场景

当前视觉导航研究存在可改进空间。首先,直接采用RNN和Transformer常忽视具身智能与传统序列建模的差异,限制了在具身智能任务中的表现。其次,依赖任务特定配置(如预训练模块和数据集特定逻辑)降低了方法的泛化性。本文从因果视角分析导航任务与其他序列任务的本质差异,提出因果感知Transformer(CAT)网络,包含因果理解模块以增强环境理解能力。该方法无任务特定归纳偏置,支持端到端训练,提升了跨场景泛化性。实证评估表明,该方法在多种设置、任务和仿真环境中均优于基准模型。消融实验显示性能提升主要归因于因果理解模块,在强化学习和监督学习中均表现出有效性与高效性。

原文摘要 · Abstract (English)

Current research in Visual Navigation reveals opportunities for improvement. First, the direct adoption of RNNs and Transformers often overlooks the specific differences between Embodied AI and traditional sequential data modelling, potentially limiting its performance in Embodied AI tasks. Second, the reliance on task-specific configurations, such as pre-trained modules and dataset-specific logic, compromises the generalizability of these methods. We address these constraints by initially exploring the unique differences between Navigation tasks and other sequential data tasks through the lens of Causality, presenting a causal framework to elucidate the inadequacies of conventional sequential methods for Navigation. By leveraging this causal perspective, we propose Causality-Aware Transformer (CAT) Networks for Navigation, featuring a Causal Understanding Module to enhance the models's Environmental Understanding capability. Meanwhile, our method is devoid of task-specific inductive biases and can be trained in an End-to-End manner, which enhances the method's generalizability across various contexts. Empirical evaluations demonstrate that our methodology consistently surpasses benchmark performances across a spectrum of settings, tasks and simulation environments. Extensive ablation studies reveal that the performance gains can be attributed to the Causal Understanding Module, which demonstrates effectiveness and efficiency in both Reinforcement Learning and Supervised Learning settings.

机器人导航因果推理Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。