arXiv:2509.24908cs.CL2025-09被引 1

构建西班牙法律文书极简摘要数据集,提升法律文本提炼效率

BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications

  • 基于西班牙官方公报构建3648条法律文书摘要数据集
  • 微调模型相比零样本模型准确率提升24%(41.6% vs 33.5%)
  • 适合法律信息化、司法AI领域研究者使用

由于信息过载,长文档的简洁摘要能力日益重要,但西班牙语文档尤其是法律领域的摘要资源严重不足。本文提出BOE-XSUM,一个包含3,648条来自西班牙《国家官方公报》(BOE)的简洁、通俗语言摘要的数据集。每条数据包含摘要、原文及文档类型标签。我们评估了在该数据集上微调的中等规模大语言模型性能,并与通用生成模型在零样本设置下对比。结果表明,微调模型显著优于未专门训练的模型。其中表现最佳的BERTIN GPT-J 6B(32位精度)模型相较最优零样本模型DeepSeek-R1,准确率提升24%(分别为41.6%和33.5%)。

原文摘要 · Abstract (English)

The ability to summarize long documents succinctly is increasingly important in daily life due to information overload, yet there is a notable lack of such summaries for Spanish documents in general, and in the legal domain in particular. In this work, we present BOE-XSUM, a curated dataset comprising 3,648 concise, plain-language summaries of documents sourced from Spain's ``Bolet\'ın Oficial del Estado'' (BOE), the State Official Gazette. Each entry in the dataset includes a short summary, the original text, and its document type label. We evaluate the performance of medium-sized large language models (LLMs) fine-tuned on BOE-XSUM, comparing them to general-purpose generative models in a zero-shot setting. Results show that fine-tuned models significantly outperform their non-specialized counterparts. Notably, the best-performing model -- BERTIN GPT-J 6B (32-bit precision) -- achieves a 24\% performance gain over the top zero-shot model, DeepSeek-R1 (accuracies of 41.6\% vs.\ 33.5\%).

法律AI文本摘要西班牙语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。