arXiv:2606.08025cs.CL2026-06中稿 · EMNLP被引 1

构建阿拉伯语多文体分句数据集,发现小模型在复杂场景下优于大模型。

Arabic Sentence Segmentation Across Genres and Punctuation Conditions

  • 构建涵盖八类文体的阿拉伯语分句数据集,覆盖多种标点与文档结构。
  • 轻量级编码器在最困难条件下表现超越大语言模型。
  • 分句准确率提升显著改善依存句法分析效果,适合自然语言处理研究者。

阿拉伯语句子分割因标点模糊不一致而困难,许多文本缺乏可靠的句尾标记。现有方法严重依赖标点信号,且通常在格式良好的文本上评估,难以适应真实场景。为此,我们提出AraSEG,一个涵盖八种文体、广泛覆盖标点和文档结构条件的阿拉伯语句子分割语料库。基于AraSEG,我们在日益复杂的分割设置下评估了大语言模型(LLMs)、轻量级编码器及基于依存解析器的模型。实验表明,在最困难条件下,轻量级编码器甚至优于依赖解析模型。我们进一步研究训练数据规模和文体多样性的影响,发现性能最终趋于饱和,跨文体泛化仍具挑战性。此外,我们证明准确的句子分割能显著提升下游依存句法分析效果。代码、数据与模型均已开源。

原文摘要 · Abstract (English)

Sentence segmentation in Arabic is challenging due to ambiguous and inconsistent punctuation, with many texts lacking reliable sentence boundary markers. Existing approaches rely heavily on punctuation cues and are typically evaluated on well-formed text, limiting their robustness in realistic Arabic settings. To address this, we introduce AraSEG, a genre-diverse sentence segmentation corpus spanning eight genres and a wide range of punctuation and document structure conditions. Using AraSEG, we evaluate LLMs, lightweight encoder models, and dependency parser-based models under increasingly challenging segmentation settings. Our experiments show that lightweight encoders, and even dependency parser-based models, outperform LLMs under the hardest conditions. We further investigate the effects of training data size and genre diversity, finding that performance eventually saturates and cross-genre generalization remains challenging. We also demonstrate that accurate sentence segmentation substantially improves downstream dependency parsing. We make our code, data, and models publicly available.

句子分割阿拉伯语自然语言处理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。