开源多语言句子简化数据集,助力低资源语言可读性研究。
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
- 五种语言统一标注,由专业人员按规范简化句子。
- 评估8个大模型发现高/低资源语言间性能差距显著。
- 适合研究多语言可读性与低资源语言文本简化者使用。
句子简化旨在降低语言复杂度以提升文本可读性,同时保持原意。然而,由于高质量数据稀缺,中低资源语言进展受限。为此,我们提出OasisSimp数据集,涵盖英语、僧伽罗语、泰米尔语、普什图语和泰语五种语言的句级简化数据。其中,泰语、普什图语和泰米尔语此前无相关数据集,僧伽罗语仅有少量数据。所有数据由受训标注员根据详细指南生成,确保语义、流畅性和语法正确性。我们在OasisSimp上评估了八个开源多语言大模型,发现高资源与低资源语言间存在显著性能差异,凸显多语言场景下的简化挑战。该数据集既是宝贵资源,也是严苛基准,揭示当前基于大模型的简化方法局限,推动低资源语言句子简化研究发展。数据集已公开:https://OasisSimpDataset.github.io/。
原文摘要 · Abstract (English)
Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource languages due to the scarcity of high-quality data. To address this gap, we introduce the OasisSimp dataset, a multilingual dataset for sentence-level simplification covering five languages: English, Sinhala, Tamil, Pashto, and Thai. Among these, no prior sentence simplification datasets exist for Thai, Pashto, and Tamil, while limited data is available for Sinhala. Each language simplification dataset was created by trained annotators who followed detailed guidelines to simplify sentences while maintaining meaning, fluency, and grammatical correctness. We evaluate eight open-weight multilingual Large Language Models (LLMs) on the OasisSimp dataset and observe substantial performance disparities between high-resource and low-resource languages, highlighting the simplification challenges in multilingual settings. The OasisSimp dataset thus provides both a valuable multilingual resource and a challenging benchmark, revealing the limitations of current LLM-based simplification methods and paving the way for future research in low-resource sentence simplification. The dataset is available at https://OasisSimpDataset.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。