用文字描述生成高效巡检路径,无需训练
Language-Guided Generation for Personalized Inspection Planning
- 基于视觉语言模型提取兴趣点并规划可观测路径
- 通过旅行商问题优化访问顺序,满足文本指令约束
- 适配航拍与水下设备,可直接用于真实场景巡检
我们提出一种无需训练的视觉语言模型(VLM)引导方法,基于文本描述高效生成目标巡检路径。不同于面向未知环境的通用视觉语言导航方法,本方法聚焦已知场景的高效巡检,广泛应用于医疗、海洋与土木工程领域。利用VLM从文本中提取兴趣点(POIs),识别既可见又符合空间约束的航点;通过与VLM交互迭代优化轨迹,保持关键区域的可见性与突出性。进一步求解旅行商问题(TSP),在满足文本隐含顺序约束的前提下,找到最优访问路径。最后通过轨迹优化生成平滑、可执行的巡检路径,适用于空中与水下机器人。我们在手工设计和真实扫描环境上进行了评估,结果表明该方法能有效生成符合用户指令的巡检路径。
原文摘要 · Abstract (English)
We propose a training-free, Vision-Language Model (VLM)-guided approach for efficiently generating trajectories to facilitate target inspection planning based on text descriptions. Unlike existing Vision-and-Language Navigation (VLN) methods designed for general agents in unknown environments, our approach specifically targets the efficient inspection of known scenes, with widespread applications in fields such as medical, marine, and civil engineering. Leveraging VLMs, our method first extracts points of interest (POIs) from the text description, then identifies a set of waypoints from which POIs are both salient and align with the spatial constraints defined in the prompt. Next, we interact with the VLM to iteratively refine the trajectory, preserving the visibility and prominence of the POIs. Further, we solve a Traveling Salesman Problem (TSP) to find the most efficient visitation order that satisfies the order constraint implied in the text description. Finally, we apply trajectory optimization to generate smooth, executable inspection paths for aerial and underwater vehicles. We have evaluated our method across a series of both handcrafted and real-world scanned environments. The results demonstrate that our approach effectively generates inspection planning trajectories that adhere to user instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。