让在线适应变成可积累的知识资产,提升视觉语言导航的跨域能力
Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation

- 用软提示和鱼骨加权机制捕获可迁移知识
- 构建动态知识库,实现跨域知识融合与持续优化
- 无需训练即可适配新环境,适合野外部署的导航系统
在非平稳环境变化下,视觉-语言导航(VLN)代理面临严峻挑战。现有测试时自适应(TTA)方法多将在线适应视为孤立、短暂的更新,导致灾难性遗忘和负迁移。为此,我们提出互域资产桥接框架(IDEA),将适应过程转化为资产的积累与组合。IDEA通过鱼骨引导加权优化软提示,捕获可迁移知识,并结合域坐标构建动态资产库。利用该库,通过将目标域投影至历史知识的凸包,构建跨域桥梁。这一设计形成互补循环:演化中的资产库支撑桥梁构建,而桥梁提供优越初始化以加速资产优化。在REVERIE、R2R和R2R-CE基准上的大量实验表明,IDEA在多个场景中均显著优于现有方法,展现出通过资产共享实现零训练适应的能力。
原文摘要 · Abstract (English)
Navigating under non-stationary environment shifts poses a critical challenge for a Vision-and-Language Navigation (VLN) agent deployed in the wild. Yet, existing Test-Time Adaptation (TTA) methods for VLN largely treat online adaptation as transient, isolated updates, leading to catastrophic forgetting and negative transfer. To overcome these issues, we propose Inter-Domain BridgE with Historical Assets (IDEA), a novel TTA framework that transforms adaptation into the accumulation and composition of assets. Specifically, IDEA introduces soft prompts optimized via a Fisher-guided weighting scheme to capture the transferable knowledge. These optimized prompts are then augmented with domain coordinates to form a dynamic asset library. Leveraging this library, IDEA constructs a cross-domain bridge by projecting the target domain onto the convex hull of historical knowledge. These designs form a complementary loop: the evolving library underpins bridge construction, while the bridge provides superior initialization to accelerate asset optimization. Extensive experiments across REVERIE, R2R, and R2R-CE benchmarks demonstrate the consistent superiority of IDEA over existing methods, showcasing its ability to enable training-free adaptation via asset sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。