构建260万条孟加拉语关键词-文本对数据集,推动低资源语言文本生成研究。
Bangla Key2Text: Text Generation from Keywords for a Low Resource Language
- 用BERT提取数百万孟加拉语新闻关键词,生成结构化数据对
- 微调mT5和BanglaT5模型显著提升关键词引导的文本生成效果
- 开源数据集与模型,助力孟加拉语自然语言生成研究
本文提出 extit{Bangla Key2Text},一个包含260万条孟加拉语关键词-文本对的大规模数据集,用于低资源语言下的关键词驱动文本生成。该数据集通过基于BERT的关键词提取管道,从数百万篇孟加拉语新闻文本中自动构建,将原始文章转化为适合监督学习的结构化数据对。为建立基准性能,我们对两种序列到序列模型 exttt{mT5} 和 exttt{BanglaT5} 进行微调,并采用多种自动评估指标与人工评价进行测试。实验结果表明,任务特定的微调相比零样本大语言模型,在孟加拉语关键词条件文本生成上表现显著提升。数据集、训练好的模型及代码已公开发布,以支持未来在孟加拉语自然语言生成和关键词转文本任务中的研究。
原文摘要 · Abstract (English)
This paper introduces \textit{Bangla Key2Text}, a large-scale dataset of $2.6$ million Bangla keyword--text pairs designed for keyword-driven text generation in a low-resource language. The dataset is constructed using a BERT-based keyword extraction pipeline applied to millions of Bangla news texts, transforming raw articles into structured keyword--text pairs suitable for supervised learning. To establish baseline performance on this new benchmark, we fine-tune two sequence-to-sequence models, \texttt{mT5} and \texttt{BanglaT5}, and evaluate them using multiple automatic metrics and human judgments. Experimental results show that task-specific fine-tuning substantially improves keyword-conditioned text generation in Bangla compared to zero-shot large language models. The dataset, trained models, and code are publicly released to support future research in Bangla natural language generation and keyword-to-text generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。