arXiv:2607.12874cs.CVq-bio.QM2026-07

用量化指标指导合成图像生成,提升模型性能。

Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI

论文配图:Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI
图 1 · 摘自论文原文
  • 引入可量化的视觉真实度指标优化合成场景渲染。
  • 合成数据量增大会显著提升零样本检测准确率。
  • 适合需要高质量合成数据的科研与智能代理系统。

科学应用中的深度学习计算机视觉需大量人工标注数据,过程繁琐且易出错。通过3D建模与渲染生成合成数据可简化流程并提高标注精度。然而,缩小真实与合成图像之间的领域差距缺乏系统性量化指导。我们提出GraNatPy,一个基于指标引导渲染场景优化的Python工具包。实验表明,提升合成数据的真实感、多样性和规模能增强视觉感知,并显著提高目标检测模型的零样本性能。此外,基于病毒斑点照片的研究发现,梯度相似性影响小目标检测表现,混合真实与合成数据可有效改善。最后,我们将程序化数据渲染转化为智能体技能(SynthClaw),实现参数自动优化。

原文摘要 · Abstract (English)

Deep learning computer vision for scientific applications requires collecting and annotating large datasets in a laborious, expensive and error-prone process. Synthetic data generation through 3D modelling and rendering may simplify this process and increase the accuracy of annotations by generating them programmatically. However, minimising the domain gap between real and synthetic images visually is subjective and lacks systematic quantitative guidance. We present GraNatPy, a Python package with metrics to guide improvement of the rendered scene. We show that quantifiable increase in realism, diversity and size of rendered dataset correlates with improved visual perception of the scene and higher zero-shot performance of an object detection model. Furthermore, we demonstrated using photographs of virological plaque assays that gradient similarity affects performance on small object detection, which can be improved by mixing real and synthetic data. Finally, we turn procedural data rendering into an agentic skill (SynthClaw) to automate the procedural parameter optimisation.

合成数据视觉检测智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。