混合真实与合成数据训练神经网络,提升模型在真实场景的泛化能力。
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
- 对比两种主流混合数据训练策略,评估其在不同架构上的表现。
- 在三种混合数据集上测试,发现合成数据占比影响模型性能与鲁棒性。
- 为提升模型真实世界适应性,提供合成数据使用优化建议,适合数据受限场景研究者。
合成数据已成为训练人工神经网络(ANN)的低成本替代方案。然而,合成数据与真实数据之间的差异导致领域差距,使训练后的模型在真实场景中表现不佳且泛化能力差。为此,研究人员提出了结合合成与真实数据的混合训练策略以缩小该差距。尽管已有一定成效,但这些策略在不同任务和网络架构下的普适性与鲁棒性仍缺乏系统评估。本研究针对三种主流神经网络架构与三个不同的混合数据集,全面分析了两种广泛使用的混合策略。通过从各数据集中采样不同比例的合成与真实数据子集,探究合成与真实数据成分对模型性能的影响。结果为优化任何ANN训练中合成数据的使用提供了关键洞见,有助于增强模型的鲁棒性与实际应用效能。
原文摘要 · Abstract (English)
Synthetic data has emerged as a cost-effective alternative to real data for training artificial neural networks (ANN). However, the disparity between synthetic and real data results in a domain gap. That gap leads to poor performance and generalization of the trained ANN when applied to real-world scenarios. Several strategies have been developed to bridge this gap, which combine synthetic and real data, known as mixed training using hybrid datasets. While these strategies have been shown to mitigate the domain gap, a systematic evaluation of their generalizability and robustness across various tasks and architectures remains underexplored. To address this challenge, our study comprehensively analyzes two widely used mixing strategies on three prevalent architectures and three distinct hybrid datasets. From these datasets, we sample subsets with varying proportions of synthetic to real data to investigate the impact of synthetic and real components. The findings of this paper provide valuable insights into optimizing the use of synthetic data in the training process of any ANN, contributing to enhancing robustness and efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。