arXiv:2411.14811cs.CVcs.CL2024-11被引 3

用贝叶斯优化生成精细视觉负样本,提升语言导航的跨模态对齐效果。

Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization

  • 基于贝叶斯优化构建对抗性视觉负样本,增强跨模态嵌入
  • 在R2R和REVERIE上实现导航性能显著提升
  • 适合研究视觉语言对齐与机器人导航的开发者

本文针对视觉-语言导航(VLN)任务中细粒度对齐的挑战,提出一种基于贝叶斯优化的对抗性优化框架,用于生成精细的对比视觉样本。现有方法依赖对比学习对齐语言与视觉轨迹,但难以处理细粒度视觉负样本。本研究通过引入贝叶斯优化生成更具挑战性的负样本,从而增强跨模态嵌入表示。在R2R和REVERIE两个主流VLN基准测试上验证了该方法的有效性,实验结果表明,改进后的嵌入能显著提升导航性能。相关代码与训练模型已公开于https://anonymous.4open.science/r/FGVLN。

原文摘要 · Abstract (English)

This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural language instructions. Current approaches use contrastive learning to align language with visual trajectory sequences. Nevertheless, they encounter difficulties with fine-grained vision negatives. To enhance cross-modal embeddings, we introduce a novel Bayesian Optimization-based adversarial optimization framework for creating fine-grained contrastive vision samples. To validate the proposed methodology, we conduct a series of experiments to assess the effectiveness of the enriched embeddings on fine-grained vision negatives. We conduct experiments on two common VLN benchmarks R2R and REVERIE, experiments on the them demonstrate that these embeddings benefit navigation, and can lead to a promising performance enhancement. Our source code and trained models are available at: https://anonymous.4open.science/r/FGVLN.

视觉语言导航贝叶斯优化对比学习跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。