arXiv:2606.22424cs.CV2026-06中稿 · ECCV

提出FlowDec,提升视觉语言导航在有图像噪声时的鲁棒性。

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation

论文配图:FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
图 1 · 摘自论文原文
  • 用时序条件流+动作中心过滤,动态修复带噪图像。
  • 导航准确率显著提升,生成延迟更低。
  • 适合真实场景下需要稳定视觉理解的任务。

连续环境中的视觉-语言导航(VLN-CE)要求智能体在未见场景中遵循自然语言指令。尽管大模型已推动该领域发展,但其性能仍受真实世界视觉退化严重影响,这一关键约束却鲜被研究。本文提出时序条件流去噪器(FlowDec),一种专为基于大模型的VLN-CE设计的图像恢复框架。FlowDec采用混合时序条件策略,使生成流程与历史上下文对齐,并通过动作中心引导的滤波机制动态评估与融合输出。大量实验表明,FlowDec在导航准确率和生成延迟方面均优于现有去噪方法。本方法建立了一种高效、鲁棒的具身导航范式,适用于不可预测的真实世界条件。

原文摘要 · Abstract (English)

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large Models (LMs) have advanced VLN-CE, their performance remains severely degraded by real-world visual corruptions, a critical yet underexplored domain constraint. We introduce Temporal Conditional Flow Decorruptor (FlowDec), a novel image restoration framework tailored for LM-based VLN-CE. FlowDec integrates a hybrid temporal conditioning strategy to align the generative flow path with historical context and employs action-centroid guided filtering to dynamically assess and integrate outputs. Extensive experiments demonstrate that FlowDec outperforms state-of-the-art decorruption methods in both navigation accuracy and generation latency. Our approach establishes a robust, efficient paradigm for resilient embodied navigation in unpredictable real-world conditions.

视觉语言导航图像去噪大模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。