揭示CUDA随机性对视觉任务可复现性的干扰,发现性能波动达4.77%
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
- 在隔离环境下测试标准与真实数据集,量化CUDA随机性影响
- 随机性可导致性能差异最高达4.77%,显著影响结果一致性
- 控制随机性虽增时但代价有限,比以往研究低估了可行性
可复现性是科学研究的基础。然而,在计算机视觉中,由于多种因素影响,获得一致结果仍具挑战性。一个关键却常被忽视的因素是CUDA引起的随机性。尽管CUDA能加速算法在GPU上的执行,但若未加以控制,其行为在多次运行中具有非确定性。尽管机器学习中的可复现性问题已被研究,但CUDA随机性在实际应用中的影响尚不明确。本研究聚焦于单一标准基准数据集及两个真实世界数据集,在隔离环境中评估该随机性的影响。结果显示,CUDA诱导的随机性可能导致性能分数差异高达4.77%。我们发现,为实现可复现性而管理这种变异性可能带来运行时间增加或性能下降,但此类代价远低于以往研究报道的程度。
原文摘要 · Abstract (English)
Reproducibility is essential for scientific research. However, in computer vision, achieving consistent results is challenging due to various factors. One influential, yet often unrecognized, factor is CUDA-induced randomness. Despite CUDA's advantages for accelerating algorithm execution on GPUs, if not controlled, its behavior across multiple executions remains non-deterministic. While reproducibility issues in ML being researched, the implications of CUDA-induced randomness in application are yet to be understood. Our investigation focuses on this randomness across one standard benchmark dataset and two real-world datasets in an isolated environment. Our results show that CUDA-induced randomness can account for differences up to 4.77% in performance scores. We find that managing this variability for reproducibility may entail increased runtime or reduce performance, but that disadvantages are not as significant as reported in previous studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。