构建多形式文本标注数据集,提升钢铁表面缺陷的视觉语言理解能力
SteelDefectX: A Multi-Form Vision-Language Dataset and Benchmark for Steel Surface Defect Analysis
- 提供三种文本形式标注:自由描述、结构化属性、模板句式
- 7778张图像覆盖25类缺陷,支持分类、分割、跨数据集迁移等任务
- 适合工业视觉语言模型研究者,尤其关注语义对齐与泛化性能
钢铁表面缺陷分析对工业质量控制至关重要,但现有基准主要依赖标签注释,限制了细粒度语义理解与视觉语言模型的系统评估。为弥补这一空白,我们提出SteelDefectX,一个包含多形式文本注释的视觉语言数据集,涵盖25类缺陷的7,778张图像。在类别层面,提供缺陷名称、代表性视觉特征及工业成因;在样本层面,每张图像配备三种文本表示:(1)自由表述的自然语言描述,(2)结构化属性标注,(3)模板化句子。这些注释提供不同表达力与可控性的文本监督。我们进一步建立涵盖视觉语言分类、分割及跨数据集迁移的综合基准,并包含检索与文本引导定位等附加评估。实验揭示文本表示中结构与灵活性的权衡:结构化属性带来更稳定的语义对齐,而自然语言描述增强迁移性与细粒度空间定位能力。结果凸显文本设计在工业视觉语言学习中的关键作用。SteelDefectX为研究工业场景下语义对齐与泛化提供了新基准。代码与数据集见https://github.com/Zhaosxian/SteelDefectX。
原文摘要 · Abstract (English)
Steel surface defect analysis is critical for industrial quality control, yet existing benchmarks rely primarily on label-only annotations, limiting fine-grained semantic understanding and systematic evaluation of vision-language models. To address this gap, we introduce SteelDefectX, a vision-language dataset with multi-form textual annotations for steel surface defect analysis, comprising 7,778 images across 25 defect categories. At the class level, the dataset provides defect names, representative visual attributes, and industrial causes. At the sample level, each image is annotated with three forms of textual representations: (1) free-form natural language descriptions, (2) structured attribute annotations, and (3) template-based sentences. These annotations provide flexible textual supervision with varying levels of expressiveness and controllability. We further establish a comprehensive benchmark covering vision-language classification, segmentation, and cross-dataset transfer, along with additional evaluations such as retrieval and text-guided localization. Experimental results reveal a trade-off between structure and flexibility in textual representations. Structured attributes provide more stable semantic alignment, while natural language descriptions improve transferability and fine-grained spatial grounding. These findings highlight the critical role of textual design in industrial vision-language learning. SteelDefectX provides a new benchmark for studying semantic alignment and generalization in industrial vision-language learning. The code and dataset are available at https://github.com/Zhaosxian/SteelDefectX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。