arXiv:2412.01293cs.CL2024-12被引 4

为僧伽罗语构建首个句级文本简化数据集,支持零资源模型评估。

SiTSE: Sinhala Text Simplification Dataset and Evaluation

  • 人工标注1000个复杂句及3000个简化句,由三名标注员完成。
  • 引入中间任务迁移学习,显著优于现有零资源方法。
  • 适用于低资源语言文本简化研究者,推动评估标准改进。

文本简化在低资源语言中研究极少,相关人工标注数据集稀缺。本文构建了首个僧伽罗语的句级文本简化数据集,包含1000个复杂句子及其由三位不同标注员生成的3000个简化版本。我们将该任务建模为多语言模型mT5和mBART上的零资源序列到序列(seq-seq)问题,利用相关任务的辅助数据,并探索中间任务迁移学习(ITTL)的可能性。实验表明,ITTL优于先前提出的零资源方法。研究还揭示了文本简化系统评估的挑战,支持开发更适合低资源语言的自动化评估指标。代码与数据已公开:https://github.com/brainsharks-fyp17/Sinhala-Text-Simplification-Dataset-and-Evaluation。

原文摘要 · Abstract (English)

Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this paper, we present a human curated sentence-level text simplification dataset for the Sinhala language. Our evaluation dataset contains 1,000 complex sentences and corresponding 3,000 simplified sentences produced by three different human annotators. We model the text simplification task as a zero-shot and zero resource sequence-to-sequence (seq-seq) task on the multilingual language models mT5 and mBART. We exploit auxiliary data from related seq-seq tasks and explore the possibility of using intermediate task transfer learning (ITTL). Our analysis shows that ITTL outperforms the previously proposed zero-resource methods for text simplification. Our findings also highlight the challenges in evaluating text simplification systems, and support the calls for improved metrics for measuring the quality of automated text simplification systems that would suit low-resource languages as well. Our code and data are publicly available: https://github.com/brainsharks-fyp17/Sinhala-Text-Simplification-Dataset-and-Evaluation

文本简化低资源语言数据集迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。