用风格迁移生成古埃及象形文字数据集,提升低资源语言模型训练效果
Neural Style Transfer for Synthesising a Dataset of Ancient Egyptian Hieroglyphs
- 通过神经风格迁移将字体转化为古埃及象形文字样式
- 生成数据训练的模型在真实未见图像上表现与真实数据相当
- 适合古文字数字化、文化遗产保护研究者使用
低资源语言的训练数据稀缺,制约机器学习应用。古埃及语即为典型例子。本文提出一种新方法:利用神经风格迁移(NST)技术,将数字字体转换为古埃及象形文字样式,从而合成训练数据集。实验表明,基于该合成数据训练的图像分类模型,在真实未见过的象形文字图像上,性能与使用真实照片训练的模型相当,且具备良好的可迁移性。
原文摘要 · Abstract (English)
The limited availability of training data for low-resource languages makes applying machine learning techniques challenging. Ancient Egyptian is one such language with few resources. However, innovative applications of data augmentation methods, such as Neural Style Transfer, could overcome these barriers. This paper presents a novel method for generating datasets of ancient Egyptian hieroglyphs by applying NST to a digital typeface. Experimental results found that image classification models trained on NST-generated examples and photographs demonstrate equal performance and transferability to real unseen images of hieroglyphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。