arXiv:2505.01881cs.CVcs.AI2025-05CVPR被引 3

融合视觉语言模型与传感器数据,实现更可靠、可解释的导航决策。

PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications

  • 双分支架构同步生成动作与链式思考解释
  • 自适应卡尔曼滤波提升20%以上导航成功率
  • 适合需要透明决策的自动驾驶与机器人场景

在多样环境与领域中实现鲁棒导航,需兼顾精确状态估计与透明决策。本文提出PhysNav-DG框架,将经典传感器融合与视觉语言模型的语义能力结合。其双分支结构从多传感器输入中预测导航动作,并同时生成详细的链式思考解释。改进的自适应卡尔曼滤波器根据环境上下文动态调整噪声参数,融合原始传感器数据及LLaMA 3.2 11B、BLIP-2等模型的语义信息。为评估该方法,我们构建了MD-NEX基准数据集,统一室内导航、自动驾驶与社交导航任务,包含真实动作标签与人工验证的解释。大量实验与消融分析表明,PhysNav-DG在导航成功率上提升超过20%,且解释内容高度可信且清晰。本工作实现了高层语义推理与几何规划的融合,推动更安全、可信的自主系统发展。

原文摘要 · Abstract (English)

Robust navigation in diverse environments and domains requires both accurate state estimation and transparent decision making. We present PhysNav-DG, a novel framework that integrates classical sensor fusion with the semantic power of vision-language models. Our dual-branch architecture predicts navigation actions from multi-sensor inputs while simultaneously generating detailed chain-of-thought explanations. A modified Adaptive Kalman Filter dynamically adjusts its noise parameters based on environmental context. It leverages several streams of raw sensor data along with semantic insights from models such as LLaMA 3.2 11B and BLIP-2. To evaluate our approach, we introduce the MD-NEX Benchmark, a novel multi-domain dataset that unifies indoor navigation, autonomous driving, and social navigation tasks with ground-truth actions and human-validated explanations. Extensive experiments and ablations show that PhysNav-DG improves navigation success rates by over 20% and achieves high efficiency, with explanations that are both highly grounded and clear. This work connects high-level semantic reasoning and geometric planning for safer and more trustworthy autonomous systems.

导航系统多模态融合可解释性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。