arXiv:2507.15365cs.CV2025-07ICCV被引 12

用高质量合成数据训练视觉模型,效率更高且精度不降。

DAViD: Data-efficient and Accurate Vision Models from Synthetic Data

  • 用程序生成的高保真合成数据训练模型
  • 在深度估计等三任务上达到与大模型相当的精度
  • 适合需要高效训练和公平性的研究者使用

当前以人为中心的计算机视觉模型虽性能优异,但通常需数十亿参数、大规模数据集和高昂算力。本文证明:仅用小规模但高保真合成数据即可训练出同样高精度且更高效的模型。合成数据提供精细细节与完美标签,并确保数据来源清晰、使用合规。通过程序化生成可精确控制数据多样性,从而缓解模型不公平性。在真实图像上对深度估计、表面法向估计和软前景分割三项密集预测任务的定量评估显示,该模型精度与同类基础模型相当,但训练与推理成本仅为后者的极小部分。所用合成数据集及训练模型已公开:https://aka.ms/DAViD。

原文摘要 · Abstract (English)

The state of the art in human-centric computer vision achieves high accuracy and robustness across a diverse range of tasks. The most effective models in this domain have billions of parameters, thus requiring extremely large datasets, expensive training regimes, and compute-intensive inference. In this paper, we demonstrate that it is possible to train models on much smaller but high-fidelity synthetic datasets, with no loss in accuracy and higher efficiency. Using synthetic training data provides us with excellent levels of detail and perfect labels, while providing strong guarantees for data provenance, usage rights, and user consent. Procedural data synthesis also provides us with explicit control on data diversity, that we can use to address unfairness in the models we train. Extensive quantitative assessment on real input images demonstrates accuracy of our models on three dense prediction tasks: depth estimation, surface normal estimation, and soft foreground segmentation. Our models require only a fraction of the cost of training and inference when compared with foundational models of similar accuracy. Our human-centric synthetic dataset and trained models are available at https://aka.ms/DAViD.

合成数据视觉模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。