arXiv:2509.09990cs.CL2025-09EMNLP被引 4

构建少数民族语言新闻标题生成数据集,助力中文少数语种研究

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

  • 收集藏语10万条、维吾尔语和蒙古语各5万条新闻标题数据
  • 设计母语者标注的高质量测试集,可作为领域基准
  • 填补中国少数民族语言自动标题生成的数据空白

中国少数民族语言如藏语、维吾尔语和传统蒙古语因书写系统与国际标准不同,面临严重语料匮乏问题,尤其在监督学习任务如新闻标题生成中。为解决此问题,我们提出了一个新数据集——中国少数民族新闻标题生成数据集(CMHG),包含藏语10万条、维吾尔语和蒙古语各5万条,专为新闻标题生成任务设计。同时,我们构建了一个由母语者标注的高质量测试集,旨在成为该领域未来研究的基准。本数据集期望推动中国少数民族语言新闻标题生成的发展,并促进相关基准的建立。

原文摘要 · Abstract (English)

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation. To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks. Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain. We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks.

新闻生成少数民族语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。