arXiv:2501.09267cs.CVcs.RO2025-01中稿 · presentation at th…被引 1

对比开集模型与微调轻量模型在工地机电设备检测中的表现

Are Open-Vocabulary Models Ready for Detection of MEP Elements on Construction Sites

  • 用移动地面机器人采集数据,对比开集视觉语言模型与微调轻量检测器
  • 微调轻量模型在专业场景下检测精度显著优于开集模型
  • 适合关注工地自动化与视觉检测落地的工程应用研究者

建筑行业长期探索机器人与计算机视觉技术,但其在工地的实际部署仍极为有限。这些技术有望通过提升施工管理的准确性、效率与安全性,彻底改变传统流程。配备先进视觉系统的地面机器人可自动监测机械、电气、管道(MEP)系统。本研究评估了开集视觉-语言模型与微调后的轻量级闭集目标检测器在移动端机器人平台上识别工地MEP组件的适用性。通过安装在地面机器人的摄像头采集数据并人工标注分析,结果表明:尽管视觉-语言模型具有更强的泛化能力,但在专业环境和特定任务中,微调的轻量模型仍显著优于前者。

原文摘要 · Abstract (English)

The construction industry has long explored robotics and computer vision, yet their deployment on construction sites remains very limited. These technologies have the potential to revolutionize traditional workflows by enhancing accuracy, efficiency, and safety in construction management. Ground robots equipped with advanced vision systems could automate tasks such as monitoring mechanical, electrical, and plumbing (MEP) systems. The present research evaluates the applicability of open-vocabulary vision-language models compared to fine-tuned, lightweight, closed-set object detectors for detecting MEP components using a mobile ground robotic platform. A dataset collected with cameras mounted on a ground robot was manually annotated and analyzed to compare model performance. The results demonstrate that, despite the versatility of vision-language models, fine-tuned lightweight models still largely outperform them in specialized environments and for domain-specific tasks.

工地检测视觉语言模型轻量检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。