arXiv:2607.20078cs.CL2026-07EMNLP被引 3

首个罗马尼亚语词汇简化数据集及系统,助力自然语言处理研究

RALS: Resources and Baselines for Romanian Automatic Lexical Simplification

  • 构建首个罗马尼亚语词汇复杂度预测与简化联合数据集
  • 提供3921个词在上下文中的人工复杂度标注
  • 提出排序简化建议的方法,适合语言技术开发者使用

我们推出了首个同时涵盖罗马尼亚语词汇复杂度预测(LCP)标注和词汇简化(LS)的数据集,并对比了多种词汇简化方法。提出一种基于成对排序近似的方法,利用独立的人工判断结果,将简化候选词按从简单到复杂的顺序排列。此外,我们提供了3,921个词在上下文中的人工词汇复杂度标注。最后,探索了几种新颖的复杂度预测与简化流程,首次构建了罗马尼亚语文本简化系统。

原文摘要 · Abstract (English)

We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, along with a comparison of lexical simplification approaches. We propose a methodology for ordering simplification suggestions using a pairwise ranking approximation method, arranging candidates from simple to complex based on a separate set of human judgments. In addition, we provide human lexical complexity annotations for 3,921 word samples in context. Finally, we explore several novel pipelines for complexity prediction and simplification and present the first text simplification system for Romanian.

词汇简化罗马尼亚语数据集NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。