arXiv:2601.01543cs.CLcs.AI2026-01

用英文数据自动生成高质量印地语摘要数据集

Bridging the Data Gap: Creating a Hindi Text Summarization Dataset from the English XSUM

  • 基于XSUM英文数据,通过翻译与语言适配构建印地语摘要集
  • 使用COMET和大模型筛选,保证摘要质量与上下文相关性
  • 为低资源语言提供可复用的自动化数据生成方法

当前自然语言处理进展主要集中在资源丰富的语言,导致像印地语这样的低资源语言在文本摘要领域缺乏高质量数据集。为弥补这一差距,本文提出一种低成本、自动化的框架,基于英文极端摘要(XSUM)数据集,结合先进的翻译与语言适应技术,生成综合性印地语文本摘要数据集。为确保高保真度与上下文相关性,采用跨语言翻译评估优化指标(COMET)进行验证,并辅以大语言模型(LLMs)进行内容筛选。最终数据集涵盖多样化主题,结构复杂度与原始XSUM相当。该工作不仅为印地语NLP研究提供直接工具,还提供了一种可扩展的方法论,推动其他未充分覆盖语言的NLP发展,降低数据构建成本,助力更贴近文化的计算语言学模型研发。

原文摘要 · Abstract (English)

Current advancements in Natural Language Processing (NLP) have largely favored resource-rich languages, leaving a significant gap in high-quality datasets for low-resource languages like Hindi. This scarcity is particularly evident in text summarization, where the development of robust models is hindered by a lack of diverse, specialized corpora. To address this disparity, this study introduces a cost-effective, automated framework for creating a comprehensive Hindi text summarization dataset. By leveraging the English Extreme Summarization (XSUM) dataset as a source, we employ advanced translation and linguistic adaptation techniques. To ensure high fidelity and contextual relevance, we utilize the Crosslingual Optimized Metric for Evaluation of Translation (COMET) for validation, supplemented by the selective use of Large Language Models (LLMs) for curation. The resulting dataset provides a diverse, multi-thematic resource that mirrors the complexity of the original XSUM corpus. This initiative not only provides a direct tool for Hindi NLP research but also offers a scalable methodology for democratizing NLP in other underserved languages. By reducing the costs associated with dataset creation, this work fosters the development of more nuanced, culturally relevant models in computational linguistics.

文本摘要低资源语言数据构建跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。