arXiv:2607.19288cs.CVcs.RO2026-07

无需训练,通过测试时迭代优化无人机导航路径

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

论文配图:No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation
图 1 · 摘自论文原文
  • 测试时迭代生成多条路径并自我修正
  • 在复杂环境中提升飞行安全与路径准确率
  • 适合追求高可靠性的无人机自主导航应用

测试时缩放为提升视觉语言模型(VLM)的推理性能提供了新思路,且无需额外训练。现有无人机(UAV)视觉语言导航(VLN)方法通常依赖单次推理,复杂环境下易产生次优或不安全轨迹。本文提出一种简单有效的方法,将测试时缩放应用于UAV VLN:通过无训练的迭代优化过程,引导模型重新评估初始规划,提升准确性与安全性。方法首先并行生成多个候选路径,再通过自校正步骤进行筛选;设计多准则评分函数,综合评估路径的安全性、目标对齐度和前进进度。该组合无需修改基础模型,使冻结的UAV导航VLM实现自我纠错,生成更可靠飞行计划,在该任务上达到当前最优(SOTA)性能。

原文摘要 · Abstract (English)

Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories. In this paper, we explore a simple and effective approach to apply test-time scaling to VLN for UAV. We enhance navigation reasoning through an iterative refinement process that requires no extra model training, guiding the model to re-evaluate its initial navigation plan for better accuracy and safety. Our method first prompts the model to generate multiple parallel candidates and then performs a self-correction step, achieving deeper and more robust planning without changing the underlying model. To further strengthen decision-making, we design a multi-criteria scoring function to evaluate the refined candidates based on safety, goal alignment, and forward-progress. This simple yet powerful combination enables a frozen UAV navigation VLMs to self-correct and generate more accurate and reliable flight plans, achieving SOTA performance in this task.

无人机导航视觉语言模型测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。