arXiv:2603.14276cs.CVcs.AI2026-03被引 4

提出TuKA框架,实现跨场景长期导航的持续学习。

All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation

  • 用高阶张量与塔克分解建模多层级导航知识
  • 在多个场景中持续学习,避免灾难性遗忘
  • 适合需要长期部署的智能导航系统

部署视觉-语言导航(VLN)代理需适应多样场景与环境,但针对特定场景微调常导致其他场景出现灾难性遗忘,严重限制其长期灵活应用。本文将此问题形式化为全天多场景终身VLN(AML-VLN)挑战。现有参数高效适配器(如LoRA及其变体)受限于二维矩阵结构,难以捕捉跨多场景、多层次的导航知识。为此,我们提出塔克适配(TuKA),将多层级导航知识表示为高阶张量,并通过塔克分解将其解耦为共享子空间与场景特异性专家。进一步引入解耦知识增量学习策略,在巩固共享子空间的同时约束特定专家,实现解耦式终身学习。基于TuKA,我们构建名为AlldayWalker的VLN代理,可在多个导航场景中持续学习,实现全天候多场景导航。大量实验表明,AlldayWalker持续优于当前最优基线。

原文摘要 · Abstract (English)

Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic forgetting in others, which severely limits flexible long-term deployment. We formalize this challenge as the all-day multi-scenes lifelong VLN (AML-VLN) problem. Existing parameter-efficient adapters (e.g., LoRA and its variants) are limited by their two-dimensional matrix form, which fails to capture the multi-hierarchical navigation knowledge spanning multiple scenes and environments. To address this, we propose Tucker Adaptation (TuKA), which represents the multi-hierarchical navigation knowledge as a high-order tensor and leverages Tucker decomposition to decouple the knowledge into shared subspaces and scenario-specific experts. We further introduce a decoupled knowledge incremental learning strategy to consolidate shared subspaces while constraining specific experts for decoupled lifelong learning. Building on TuKA, we also develop a VLN agent named AlldayWalker, which continually learns across multiple navigation scenarios, achieving all-day multi-scenes navigation. Extensive experiments show that AlldayWalker consistently outperforms state-of-the-art baselines.

视觉语言导航终身学习张量分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。