通过隐空间对齐,让视觉语言导航更稳定高效
LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation
- 用专家策略的隐空间约束感知与控制接口
- 在新环境、光照下零样本泛化表现优异
- 轻量适配器适合实时部署,跨模态复用性强
我们提出LCLA(语言条件隐空间对齐),一种视觉-语言导航框架,通过将感官观测对齐到专家策略的隐空间,学习模块化的感知-动作接口。专家先利用特权状态信息训练,生成足够用于控制的隐空间,随后冻结其隐接口和动作头。一个轻量级适配器被训练,将原始视觉-语言观测通过冻结的视觉-语言模型映射至专家的隐空间,将视觉运动学习问题转化为监督隐空间对齐,而非端到端策略优化。这种解耦强化了感知与控制间的稳定契约,使专家行为可跨感知模态和环境变化复用。我们在视觉-语言室内导航任务上验证LCLA,结果显示对齐的隐空间在分布内表现强,且在未见过的环境、光照条件和视角下具备鲁棒的零样本泛化能力,推理时仍保持轻量。
原文摘要 · Abstract (English)
We propose LCLA (Language-Conditioned Latent Alignment), a framework for vision-language navigation that learns modular perception-action interfaces by aligning sensory observations to a latent representation of an expert policy. The expert is first trained with privileged state information, inducing a latent space sufficient for control, after which its latent interface and action head are frozen. A lightweight adapter is then trained to map raw visual-language observations, via a frozen vision-language model, into the expert's latent space, reducing the problem of visuomotor learning to supervised latent alignment rather than end-to-end policy optimization. This decoupling enforces a stable contract between perception and control, enabling expert behavior to be reused across sensing modalities and environmental variations. We instantiate LCLA and evaluate it on a vision-language indoor navigation task, where aligned latent spaces yield strong in-distribution performance and robust zero-shot generalization to unseen environments, lighting conditions, and viewpoints while remaining lightweight at inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。