通过在线分离前景与背景,提升视觉语言导航的泛化能力。
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
- 用语义增强地标识别分离前景与背景特征。
- 在REVERIE和R2R上实现最优性能,显著提升泛化性。
- 适合研究视觉语言导航与多模态表征学习的学者。
遵循语言指令,视觉语言导航(VLN)智能体需在未见过的环境中导航。尽管多维度视觉表征增强推动了VLN进展,但视觉观测中前景与背景的作用仍被忽视。直观上,前景提供语义线索,背景蕴含空间连通信息。受此启发,我们提出共识驱动的在线特征增强策略(COFA),通过交替使用前景与背景特征,促进可导航泛化。首先,利用语义增强的地标识别,将前景与背景作为候选增强特征分离。随后,共识驱动的在线增强策略促使智能体根据多样指令和导航位置,整合两阶段投票结果以确定特征偏好。在REVERIE和R2R数据集上的实验表明,该在线前景-背景增强策略提升了基线模型的泛化能力,并达到当前最优性能。
原文摘要 · Abstract (English)
Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has propelled advancements in VLN, the significance of foreground and background in visual observations remains underexplored. Intuitively, foreground regions provide semantic cues, whereas the background encompasses spatial connectivity information. Inspired on this insight, we propose a Consensus-driven Online Feature Augmentation strategy (COFA) with alternative foreground and background features to facilitate the navigable generalization. Specifically, we first leverage semantically-enhanced landmark identification to disentangle foreground and background as candidate augmented features. Subsequently, a consensus-driven online augmentation strategy encourages the agent to consolidate two-stage voting results on feature preferences according to diverse instructions and navigational locations. Experiments on REVERIE and R2R demonstrate that our online foreground-background augmentation boosts the generalization of baseline and attains state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。