无需训练,通过测试时迭代优化无人机导航路径
No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

- 测试时迭代生成多条路径并自我修正
- 在复杂环境中提升飞行安全与路径准确率
- 适合追求高可靠性的无人机自主导航应用
测试时缩放为提升视觉语言模型(VLM)的推理性能提供了新思路,且无需额外训练。现有无人机(UAV)视觉语言导航(VLN)方法通常依赖单次推理,复杂环境下易产生次优或不安全轨迹。本文提出一种简单有效的方法,将测试时缩放应用于UAV VLN:通过无训练的迭代优化过程,引导模型重新评估初始规划,提升准确性与安全性。方法首先并行生成多个候选路径,再通过自校正步骤进行筛选;设计多准则评分函数,综合评估路径的安全性、目标对齐度和前进进度。该组合无需修改基础模型,使冻结的UAV导航VLM实现自我纠错,生成更可靠飞行计划,在该任务上达到当前最优(SOTA)性能。
原文摘要 · Abstract (English)
Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories. In this paper, we explore a simple and effective approach to apply test-time scaling to VLN for UAV. We enhance navigation reasoning through an iterative refinement process that requires no extra model training, guiding the model to re-evaluate its initial navigation plan for better accuracy and safety. Our method first prompts the model to generate multiple parallel candidates and then performs a self-correction step, achieving deeper and more robust planning without changing the underlying model. To further strengthen decision-making, we design a multi-criteria scoring function to evaluate the refined candidates based on safety, goal alignment, and forward-progress. This simple yet powerful combination enables a frozen UAV navigation VLMs to self-correct and generate more accurate and reliable flight plans, achieving SOTA performance in this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。