arXiv:2509.14755cs.CV2025-09

用扩散模型生成合成数据,提升历史画作中嗅觉相关物的检测准确率。

Data Augmentation via Latent Diffusion Models for Detecting Smell-Related Objects in Historical Artworks

  • 基于扩散模型生成合成图像,缓解标注稀疏与类别不平衡问题。
  • 加入合成数据后检测性能显著提升,小样本下效果仍稳定。
  • 适合标注成本高、数据稀缺的文物图像识别任务。

在历史艺术品中发现嗅觉相关描述是一项挑战性任务。除风格差异等作品特异性问题外,其识别需极为细致的标注类别,导致标注稀疏和极端类别不平衡。本文探索利用合成数据生成缓解上述问题,并实现对嗅觉相关物体的准确检测。我们评估了多种基于扩散模型的增强策略,结果表明将合成数据融入模型训练可有效提升检测性能。研究发现,利用扩散模型的大规模预训练能力,是一种提升检测精度的有前景方法,尤其适用于标注稀缺且获取成本高的小众应用场景。此外,该方法在数据量相对较小的情况下依然有效,进一步扩展具有巨大提升潜力。

原文摘要 · Abstract (English)

Finding smell references in historic artworks is a challenging problem. Beyond artwork-specific challenges such as stylistic variations, their recognition demands exceptionally detailed annotation classes, resulting in annotation sparsity and extreme class imbalance. In this work, we explore the potential of synthetic data generation to alleviate these issues and enable accurate detection of smell-related objects. We evaluate several diffusion-based augmentation strategies and demonstrate that incorporating synthetic data into model training can improve detection performance. Our findings suggest that leveraging the large-scale pretraining of diffusion models offers a promising approach for improving detection accuracy, particularly in niche applications where annotations are scarce and costly to obtain. Furthermore, the proposed approach proves to be effective even with relatively small amounts of data, and scaling it up provides high potential for further enhancements.

图像检测扩散模型艺术分析数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。