构建大规模基建缺陷数据集,暴露视觉大模型在真实场景中的严重短板。
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

- 构建15万张高分辨率图像的基建缺陷数据集,联合土木专家五年精标。
- 现有模型在真实基建场景下零样本性能仅达约25% mAP,专业模型亦难突破。
- 揭示当前视觉大模型在纹理缺失、形状依赖等真实场景中的根本缺陷。
自动化结构健康监测对防止基础设施灾难性失效至关重要。精确的像素级缺陷分割是评估结构完整性所必需的,但民用基础设施缺陷分割的发展因数据极度稀缺而受阻,且标注需昂贵专家参与。问题本身的算法挑战也加剧了这一困境,包括中心偏倚,以及在几乎无纹理的建筑材料上更依赖形状信息。为突破瓶颈,我们引入「基础模型中的裂缝」(Cracks in the Foundation, CiF),这是迄今为止最大、最详细的民用基础设施(实例)分割数据集,包含约15万张由土木工程专家历时五年精心标注的高分辨率图像。借助这一前所未有的数据资源,我们揭示了当前视觉AI的一个盲区:尽管可提示的基础模型(FMs)和视觉语言模型(VLMs)已取得进展,且当前专用分割模型表现优异,但在真实基建场景中实现密集图像理解仍远未解决。评估表明,即使最先进的零样本基础模型在实际基建部署中仍面临显著挑战,而经过领域特定监督的专用模型性能也停滞在约25% mAP。CiF将基建检测这一看似简单的感知任务确立为开放挑战,揭示了当前主要基于互联网图像训练的模型在理论与实践上的根本弱点。
原文摘要 · Abstract (English)
Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed to accurately assess structural integrity, but progress in defect segmentation for civil infrastructures has been held back by an extreme scarcity of data, which requires costly expert annotation. The need for data is accentuated by algorithmic hurdles intrinsic to the problem, including center-bias and the need to rely more on shape when inspecting nearly textureless building materials. To remove the bottleneck, we introduce Cracks in the Foundation (CiF), the largest and most detailed civil infrastructure (instance) segmentation dataset to date, comprising $\approx$150,000 high-resolution images meticulously curated over five years in collaboration with civil engineering experts. With the help of this unprecedented data source, we expose a blind spot of current visual AI: despite the advent of promptable Foundation Models (FMs) and Vision Language Models (VLMs), and despite the impressive abilities of today's specialised segmentation models, it turns out that dense image understanding in the built environment is nowhere near solved. Our evaluations indicate that even the most recent zero-shot FMs face significant challenges when deployed on real-world infrastructure and even the performance of specialised models with domain-specific supervision plateaus at $\approx$25% mAP. CiF establishes inspection of civil infrastructure, an elementary and seemingly easy perceptual task, as an open challenge that reveals fundamental weaknesses of present-day models trained predominantly on internet images, literally and figuratively highlighting cracks in the current foundation model paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。