用语言引导的机器人系统自动检查运河涵洞,省去人工巡检难题。
Language-in-the-Loop Culvert Inspection on the Erie Canal
- 通过语言提示让视觉模型提出检测区域并说明理由
- 在真实涵洞中实现自主移动拍摄,最终评估准确率达80%
- 无需领域微调,适合水利设施巡检人员使用
埃里运河等早期建设的运河涵洞需频繁检查以确保运行安全。由于年久失修、结构复杂、照明差、天气影响及通行困难,人工巡检极具挑战。本文提出VISION系统,一种端到端的语言引导自主巡检框架,将网络规模的视觉-语言模型(VLM)与受限视角规划相结合。通过简短指令向VLM请求开放词汇的感兴趣区域(ROI)建议及其置信度和理由,融合立体深度恢复尺度信息,并由考虑涵洞约束的规划器驱动四足机器人重新定位,获取目标细节图像。该系统在埃里运河下方涵洞中部署,实现了机载闭环‘观察-决策-移动-重成像’,生成高分辨率图像用于详细报告,且无需领域特定微调。外部评估显示,初始ROI建议与专家意见一致率达61.4%,最终重成像评估达80%,表明VISION能将初步假设转化为符合专家判断的可靠发现。
原文摘要 · Abstract (English)
Culverts on canals such as the Erie Canal, built originally in 1825, require frequent inspections to ensure safe operation. Human inspection of culverts is challenging due to age, geometry, poor illumination, weather, and lack of easy access. We introduce VISION, an end-to-end, language-in-the-loop autonomy system that couples a web-scale vision-language model (VLM) with constrained viewpoint planning for autonomous inspection of culverts. Brief prompts to the VLM solicit open-vocabulary ROI proposals with rationales and confidences, stereo depth is fused to recover scale, and a planner -- aware of culvert constraints -- commands repositioning moves to capture targeted close-ups. Deployed on a quadruped in a culvert under the Erie Canal, VISION closes the see, decide, move, re-image loop on-board and produces high-resolution images for detailed reporting without domain-specific fine-tuning. In an external evaluation by New York Canal Corporation personnel, initial ROI proposals achieved 61.4\% agreement with subject-matter experts, and final post-re-imaging assessments reached 80\%, indicating that VISION converts tentative hypotheses into grounded, expert-aligned findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。