arXiv:2512.21694cs.CVcs.AI2025-12中稿 · publication in 202…

用GAN生成多样化的孟加拉手写文字,解决数据稀缺难题

BeHGAN: Bengali Handwritten Word Generation from Plain Text Using Generative Adversarial Networks

  • 基于生成对抗网络,从文本生成孟加拉手写体
  • 自建500人参与的高质量孟加拉手写数据集
  • 适合关注低资源语言生成与手写仿真研究者

手写文本识别(HTR)是成熟的研究领域,而手写文本生成(HTG)仍处于发展初期且潜力巨大。由于个人书写风格差异显著,生成真实手写文本面临挑战,需要大规模多样化的数据集,但这类数据难以获取且不公开。孟加拉语是全球第五大使用语言,尽管英语和阿拉伯语已有相关研究,孟加拉手写文本生成却鲜少被关注。为此,本文提出一种生成孟加拉手写单词的方法。我们自建了一个包含约500名不同年龄与性别参与者的手写样本数据集,并对所有图像进行预处理以确保一致性和质量。实验表明,该方法能从输入纯文本生成多样化的手写输出。本工作推动了孟加拉手写生成技术的发展,可为后续研究提供支持。

原文摘要 · Abstract (English)

Handwritten Text Recognition (HTR) is a well-established research area. In contrast, Handwritten Text Generation (HTG) is an emerging field with significant potential. This task is challenging due to the variation in individual handwriting styles. A large and diverse dataset is required to generate realistic handwritten text. However, such datasets are difficult to collect and are not readily available. Bengali is the fifth most spoken language in the world. While several studies exist for languages such as English and Arabic, Bengali handwritten text generation has received little attention. To address this gap, we propose a method for generating Bengali handwritten words. We developed and used a self-collected dataset of Bengali handwriting samples. The dataset includes contributions from approximately five hundred individuals across different ages and genders. All images were pre-processed to ensure consistency and quality. Our approach demonstrates the ability to produce diverse handwritten outputs from input plain text. We believe this work contributes to the advancement of Bengali handwriting generation and can support further research in this area.

手写生成GAN低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。