arXiv:2409.17045cs.CVcs.AI2024-09被引 3

构建自行车工程图数据集,用大模型自动标注几何与文本信息。

GeoBiked: A Dataset with Geometric Features and Automated Labeling Techniques to Enable Deep Generative Models in Engineering Design

  • 用图像生成模型提取的超特征(Hyperfeatures)定位结构图中的几何点。
  • 多源图像标注可提升未见样本的几何点检测准确率,最高达90%以上。
  • 结合图像与类别标签能平衡文本描述的多样性与准确性,适合工程设计应用。

本文构建了包含4,355张自行车图像的GeoBiked数据集,标注了结构与技术特征,旨在推动深度生成模型在工程设计中的应用。提出两种自动化标注方法:利用图像生成模型的统一潜在特征(Hyperfeatures)检测结构图像中的几何对应关系(如轮心位置);使用GPT-4o等视觉语言模型(VLM)生成多样化的结构图像文本描述。通过将技术图像表示为Diffusion-Hyperfeatures,可实现跨图像的几何对应关系对齐。在未见样本上,引入多个已标注源图像可显著提升几何点检测精度。GPT-4o具备生成技术图像准确描述的能力,仅基于图像输入会产生幻觉,仅依赖类别标签则限制多样性。联合输入图像与标签可兼顾创意与准确性。结果表明,使用Hyperfeatures进行几何对应分析具有通用性,适用于各类技术图像的点检测与标注任务。同时,借助大模型进行文本标注虽可行,但高度依赖模型性能、提示工程与输入信息选择。本工作填补了基础模型在工程设计领域应用的研究空白,提供可用于训练、微调和条件生成的基准数据集。

原文摘要 · Abstract (English)

We provide a dataset for enabling Deep Generative Models (DGMs) in engineering design and propose methods to automate data labeling by utilizing large-scale foundation models. GeoBiked is curated to contain 4 355 bicycle images, annotated with structural and technical features and is used to investigate two automated labeling techniques: The utilization of consolidated latent features (Hyperfeatures) from image-generation models to detect geometric correspondences (e.g. the position of the wheel center) in structural images and the generation of diverse text descriptions for structural images. GPT-4o, a vision-language-model (VLM), is instructed to analyze images and produce diverse descriptions aligned with the system-prompt. By representing technical images as Diffusion-Hyperfeatures, drawing geometric correspondences between them is possible. The detection accuracy of geometric points in unseen samples is improved by presenting multiple annotated source images. GPT-4o has sufficient capabilities to generate accurate descriptions of technical images. Grounding the generation only on images leads to diverse descriptions but causes hallucinations, while grounding it on categorical labels restricts the diversity. Using both as input balances creativity and accuracy. Successfully using Hyperfeatures for geometric correspondence suggests that this approach can be used for general point-detection and annotation tasks in technical images. Labeling such images with text descriptions using VLMs is possible, but dependent on the models detection capabilities, careful prompt-engineering and the selection of input information. Applying foundation models in engineering design is largely unexplored. We aim to bridge this gap with a dataset to explore training, finetuning and conditioning DGMs in this field and suggesting approaches to bootstrap foundation models to process technical images.

工程设计生成模型图像标注大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。