arXiv:2602.07629cs.RO2026-02

通过隐空间对齐,让视觉语言导航更稳定高效

LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation

  • 用专家策略的隐空间约束感知与控制接口
  • 在新环境、光照下零样本泛化表现优异
  • 轻量适配器适合实时部署,跨模态复用性强

我们提出LCLA(语言条件隐空间对齐),一种视觉-语言导航框架,通过将感官观测对齐到专家策略的隐空间,学习模块化的感知-动作接口。专家先利用特权状态信息训练,生成足够用于控制的隐空间,随后冻结其隐接口和动作头。一个轻量级适配器被训练,将原始视觉-语言观测通过冻结的视觉-语言模型映射至专家的隐空间,将视觉运动学习问题转化为监督隐空间对齐,而非端到端策略优化。这种解耦强化了感知与控制间的稳定契约,使专家行为可跨感知模态和环境变化复用。我们在视觉-语言室内导航任务上验证LCLA,结果显示对齐的隐空间在分布内表现强,且在未见过的环境、光照条件和视角下具备鲁棒的零样本泛化能力,推理时仍保持轻量。

原文摘要 · Abstract (English)

We propose LCLA (Language-Conditioned Latent Alignment), a framework for vision-language navigation that learns modular perception-action interfaces by aligning sensory observations to a latent representation of an expert policy. The expert is first trained with privileged state information, inducing a latent space sufficient for control, after which its latent interface and action head are frozen. A lightweight adapter is then trained to map raw visual-language observations, via a frozen vision-language model, into the expert's latent space, reducing the problem of visuomotor learning to supervised latent alignment rather than end-to-end policy optimization. This decoupling enforces a stable contract between perception and control, enabling expert behavior to be reused across sensing modalities and environmental variations. We instantiate LCLA and evaluate it on a vision-language indoor navigation task, where aligned latent spaces yield strong in-distribution performance and robust zero-shot generalization to unseen environments, lighting conditions, and viewpoints while remaining lightweight at inference time.

视觉导航隐空间对齐零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。