arXiv:2412.16108cs.CVcs.AI2024-12被引 4

用ChatGPT-4 Vision看工地进度,识别能力强但定位不准

Demystifying the Potential of ChatGPT-4 Vision for Construction Progress Monitoring

  • 用高分辨率航拍图让模型分析工地进展
  • 能准确识别施工阶段、材料和机械
  • 适合想快速评估工地的工程管理者

大型视觉语言模型(LVLMs)如OpenAI的GPT-4 Vision在人工智能领域引发重要变革,尤其在视觉数据解析方面。本文研究GPT-4 Vision在建筑行业的实际应用,聚焦其对建设项目进度的监控与追踪能力。通过使用建筑工地的高分辨率航拍图像,研究分析了GPT-4 Vision在场景细节识别和时间变化追踪方面的表现。结果显示,该模型在识别施工阶段、建筑材料和机械设备方面表现良好,但在精确目标定位和分割方面存在不足。尽管如此,该技术未来仍有巨大提升空间。本研究不仅揭示了当前LVLM在建筑领域的应用现状与机遇,还探讨了通过领域特定训练、结合其他计算机视觉技术及数字孪生系统来增强模型实用性的未来方向。

原文摘要 · Abstract (English)

The integration of Large Vision-Language Models (LVLMs) such as OpenAI's GPT-4 Vision into various sectors has marked a significant evolution in the field of artificial intelligence, particularly in the analysis and interpretation of visual data. This paper explores the practical application of GPT-4 Vision in the construction industry, focusing on its capabilities in monitoring and tracking the progress of construction projects. Utilizing high-resolution aerial imagery of construction sites, the study examines how GPT-4 Vision performs detailed scene analysis and tracks developmental changes over time. The findings demonstrate that while GPT-4 Vision is proficient in identifying construction stages, materials, and machinery, it faces challenges with precise object localization and segmentation. Despite these limitations, the potential for future advancements in this technology is considerable. This research not only highlights the current state and opportunities of using LVLMs in construction but also discusses future directions for enhancing the model's utility through domain-specific training and integration with other computer vision techniques and digital twins.

视觉语言模型工地监控GPT-4

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。