arXiv:2508.21732cs.CVcs.AI2025-08被引 3

用合成数据提升大模型读取数字仪表盘的能力。

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

  • 基于3D CAD模型生成带标注的数字仪表盘图像。
  • 在真实场景测试中,模型准确率提升200%。
  • 适合做增强现实和头戴设备视觉理解的研究者。

近年来,大型视觉语言模型(LVLMs)在多模态任务中表现出色,但在真实场景下读取数字测量设备(DMDs)时仍面临挑战,如杂乱背景、遮挡、极端视角和运动模糊,常见于头戴相机和增强现实(AR)应用。为此,本文提出CAD2DMD-SET,一个用于生成合成数据的工具,通过3D CAD模型、高级渲染与高保真图像合成,构建多样化的、带VQA标签的合成DMD数据集,用于微调LVLMs。同时,我们构建了DMDBench,一个包含1,000张真实世界图像的标注验证集,用于评估模型在实际条件下的表现。使用平均归一化莱文斯坦相似度(ANLS)对三个主流LVLM进行基准测试,并结合CAD2DMD-SET生成的数据对模型进行LoRA微调,结果显示性能显著提升,InternVL模型得分提高200%,且未影响其他任务表现。这表明该合成数据集能有效提升模型在复杂条件下的鲁棒性与准确性。最终版本发布后,CAD2DMD-SET将开源,供社区扩展设备类型并自动生成数据。

原文摘要 · Abstract (English)

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities across various multimodal tasks. They continue, however, to struggle with trivial scenarios such as reading values from Digital Measurement Devices (DMDs), particularly in real-world conditions involving clutter, occlusions, extreme viewpoints, and motion blur; common in head-mounted cameras and Augmented Reality (AR) applications. Motivated by these limitations, this work introduces CAD2DMD-SET, a synthetic data generation tool designed to support visual question answering (VQA) tasks involving DMDs. By leveraging 3D CAD models, advanced rendering, and high-fidelity image composition, our tool produces diverse, VQA-labelled synthetic DMD datasets suitable for fine-tuning LVLMs. Additionally, we present DMDBench, a curated validation set of 1,000 annotated real-world images designed to evaluate model performance under practical constraints. Benchmarking three state-of-the-art LVLMs using Average Normalised Levenshtein Similarity (ANLS) and further fine-tuning LoRA's of these models with CAD2DMD-SET's generated dataset yielded substantial improvements, with InternVL showcasing a score increase of 200% without degrading on other tasks. This demonstrates that the CAD2DMD-SET training dataset substantially improves the robustness and performance of LVLMs when operating under the previously stated challenging conditions. The CAD2DMD-SET tool is expected to be released as open-source once the final version of this manuscript is prepared, allowing the community to add different measurement devices and generate their own datasets.

视觉语言模型合成数据数字仪表盘增强现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。