arXiv:2601.14791cs.CVcs.LG2026-01

用AI生成瓷器图像,提升小样本分类效果

Synthetic Data Augmentation for Multi-Task Chinese Porcelain Classification: A Stable Diffusion Approach

  • 用Stable Diffusion+LoRA生成合成瓷器图,与真实数据混合训练
  • 类型识别准确率提升5.5%,朝代和窑口任务也有小幅改善
  • 适合考古数据稀缺场景,提醒注意生成图像的真实性

考古文物分类中训练数据稀缺是深度学习应用的核心挑战,尤其针对稀有中国瓷器类型。本研究探究通过稳定扩散模型结合低秩适配(LoRA)生成的合成图像,能否有效增强有限真实数据集,在基于CNN的多任务瓷器分类中发挥作用。采用迁移学习的MobileNetV3模型,在四种分类任务(朝代、釉色、窑口、类型)上对比纯真实数据与95:5、90:10比例的真实-合成数据混合训练的效果。结果表明:类型识别提升最显著(90:10时F1-macro提高5.5%),朝代与窑口任务亦有3-4%小幅增益,说明合成数据有效性取决于生成特征与任务相关视觉特征的匹配度。本研究为生成式AI在考古研究中的应用提供实践指南,揭示了合成数据在兼顾文物真实性与数据多样性时的潜力与局限。

原文摘要 · Abstract (English)

The scarcity of training data presents a fundamental challenge in applying deep learning to archaeological artifact classification, particularly for the rare types of Chinese porcelain. This study investigates whether synthetic images generated through Stable Diffusion with Low-Rank Adaptation (LoRA) can effectively augment limited real datasets for multi-task CNN-based porcelain classification. Using MobileNetV3 with transfer learning, we conducted controlled experiments comparing models trained on pure real data against those trained on mixed real-synthetic datasets (95:5 and 90:10 ratios) across four classification tasks: dynasty, glaze, kiln and type identification. Results demonstrate task-specific benefits: type classification showed the most substantial improvement (5.5\% F1-macro increase with 90:10 ratio), while dynasty and kiln tasks exhibited modest gains (3-4\%), suggesting that synthetic augmentation effectiveness depends on the alignment between generated features and task-relevant visual signatures. Our work contributes practical guidelines for deploying generative AI in archaeological research, demonstrating both the potential and limitations of synthetic data when archaeological authenticity must be balanced with data diversity.

生成图像多任务学习考古数据Stable Diffusion

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。