arXiv:2510.06596cs.CVcs.AI2025-10中稿 · and Published at S…被引 3

提出SDQM评估生成数据质量,无需训练即可预测检测模型性能。

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

  • 基于统计特征与分布差异设计无训练评估指标
  • 与YOLO11的mAP相关性达0.92,显著优于已有方法
  • 适合资源受限场景下快速筛选高质量合成数据

机器学习模型性能高度依赖训练数据。大规模、高质量标注数据的稀缺性给构建鲁棒模型带来挑战。为此,通过仿真和生成模型产生的合成数据成为可行方案,可提升数据多样性并增强模型性能、可靠性与抗干扰能力。然而,评估此类数据质量仍需有效指标。本文提出合成数据质量度量(SDQM),可在不进行模型训练收敛的前提下评估目标检测任务的数据质量。实验表明,SDQM与YOLO11模型的mAP得分呈现强相关性(相关系数0.92),而现有指标仅表现中等或弱相关。该指标还能提供改进数据质量的可操作建议,减少昂贵的迭代训练需求。其高效可扩展特性为合成数据评估树立了新标准。代码已开源:https://github.com/ayushzenith/SDQM

原文摘要 · Abstract (English)

The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses significant challenges in creating robust models. To address this, synthetic data generated through simulations and generative models has emerged as a promising solution, enhancing dataset diversity and improving the performance, reliability, and resilience of models. However, evaluating the quality of this generated data requires an effective metric. We introduce the Synthetic Dataset Quality Metric (SDQM) to assess data quality for object detection tasks without requiring model training to converge. This metric enables more efficient generation and selection of synthetic datasets, addressing a key challenge in resource-constrained object detection tasks. In our experiments, SDQM demonstrated a strong correlation with the mean average precision (mAP) scores of YOLO11, a leading object detection model, whereas previous metrics only exhibited moderate or weak correlations. In addition, it provides actionable insights into improving dataset quality, minimizing the need for costly iterative training. This scalable and efficient metric sets a new standard for evaluating synthetic data. The code for SDQM is available at https://github.com/ayushzenith/SDQM

数据质量目标检测合成数据评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。