提出SADGE评估合成数据对真实任务的适用性,无需训练模型即可预测性能。
SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

- 融合外观与结构相似性,通过非线性交互评估合成数据质量。
- 在79,000张图像对上验证,相关性达Pearson r=0.88,显著优于单一指标。
- 适合需快速筛选高质量合成数据的研究者,尤其在标注成本高的场景。
我们提出SADGE,一种无需下游模型训练即可预测合成图像数据集在常见计算机视觉任务中表现的量化相似性度量。当前评估合成数据是否适用于真实世界任务仍是模型开发的瓶颈。现有度量(如PSNR、FID、CLIP)主要衡量真实与合成图像间的语义对齐(外观相似性),较少关注图像间几何一致性(结构相似性)。然而,据我们所知,尚无研究系统评估哪种相似性度量能最好地预测特定合成数据集的下游性能。本文在多种合成数据集和下游任务中表明,仅靠外观或结构相似性均无法可靠预测性能;真正决定合成数据效用的是二者非线性相互作用。我们测量了常用外观与几何相似性度量在物体检测、语义分割和姿态估计任务中的相关性。在五个公开的合成到真实基准家族及15种数据集变体(共79,000张图像对)上,SADGE在线性和秩相关标准下均达到最强关联,分别实现Pearson r=0.88和Spearman rho=0.77。通过组合不同几何与外观方法计算所有基准家族的SADGE得分,最优配置为融合DINOv3外观相似性与MASt3R几何一致性,通过受限双线性交互,优于最强的纯几何或纯外观基线。
原文摘要 · Abstract (English)
We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without downstream model training. Estimating whether a synthetic dataset will lead to a model that performs well on real-world data remains a bottleneck in model development. Existing evaluation metrics (e.g., PSNR, FID, CLIP) primarily measure semantic alignment between real and synthetic images (Appearance Similarity Score). Less commonly, structural similarity between images is considered to assess the domain gap (Geometric Similarity Score). However, to the best of our knowledge there exists no studies that evaluate which similarity metric is the best downstream predictor for a given synthetic dataset. In this paper, we show over a wide variety of different synthetic datasets and downstream tasks that neither appearance nor geometry alone can reliably predict downstream performance; rather, it is their non-linear interplay that dictates synthetic data utility. Specifically, we measure how commonly used Appearance and Geometric Similarity metrics computed between synthetic and real images correlate with downstream performance in object detection, semantic segmentation, and pose estimation. Across five public synthetic-to-real benchmark families and 15 dataset-level variants (79k image pairs), SADGE achieves the strongest association with downstream transfer performance under both linear and rank-based criteria, reaching Pearson r=0.88 and Spearman rho=0.77. We compute for each combination of geometry-based methods and appearance-based approaches SADGE scores across all benchmark families. The best configuration is obtained by fusing DINOv3 appearance similarity with MASt3R geometric consistency through a constrained bilinear interaction, outperforming both the strongest geometry-only baseline and the strongest appearance-only baseline .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。