arXiv:2506.05927cs.CLcs.AI2025-06

构建西班牙语行政文本简化数据集,助力自动简化系统评估。

LengClaro2023: A Dataset of Administrative Texts in Spanish with Plain Language adaptations

  • 基于西班牙社保网站高频流程,生成两种简化版本。
  • 对比不同简化策略,验证可读性提升效果。
  • 适合语言处理、无障碍信息研究者使用。

本文提出LengClaro2023,一个西班牙语法律行政文本数据集。基于西班牙社保网站最常用流程,为每篇原文创建两个简化版本:第一版遵循arText claro建议;第二版进一步采纳通用简明语言指南,探索系统优化潜力。该语言资源可用于评估西班牙语自动文本简化(ATS)系统的性能。

原文摘要 · Abstract (English)

In this work, we present LengClaro2023, a dataset of legal-administrative texts in Spanish. Based on the most frequently used procedures from the Spanish Social Security website, we have created for each text two simplified equivalents. The first version follows the recommendations provided by arText claro. The second version incorporates additional recommendations from plain language guidelines to explore further potential improvements in the system. The linguistic resource created in this work can be used for evaluating automatic text simplification (ATS) systems in Spanish.

文本简化西班牙语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。