arXiv:2503.08803cs.CL2025-03被引 1

构建首个西班牙语多文体因果关系推理数据集,助力机器理解文本因果

ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships

  • 基于BERT模型构建多文体西班牙语因果推理数据集
  • 多文体设计使模型泛化能力提升显著
  • 适合关注低资源语言与因果推理的研究者

自然语言推断(NLI),又称文本蕴涵识别(RTE),是自然语言处理中的关键领域,旨在让机器识别文本间的语义关系。尽管英语相关研究已相当充分,但针对西班牙语的工作仍相对匮乏。为此,本文提出一个多文体西班牙语NLI数据集ESNLIR,特别关注因果关系。通过引入多种文本类型,提升了模型的泛化能力。我们构建了初步基线,并使用BERT系列模型进行评估。实验结果表明,多文体数据显著增强了模型的泛化性能。相关代码、笔记本和完整数据集可在Zenodo获取:https://zenodo.org/records/15002575。仅需数据集者可访问:https://zenodo.org/records/15002371。

原文摘要 · Abstract (English)

Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), serves as a crucial area within the domain of Natural Language Processing (NLP). This area fundamentally empowers machines to discern semantic relationships between assorted sections of text. Even though considerable work has been executed for the English language, it has been observed that efforts for the Spanish language are relatively sparse. Keeping this in view, this paper focuses on generating a multi-genre Spanish dataset for NLI, ESNLIR, particularly accounting for causal Relationships. A preliminary baseline has been conceptualized and subjected to an evaluation, leveraging models drawn from the BERT family. The findings signify that the enrichment of genres essentially contributes to the enrichment of the model's capability to generalize. The code, notebooks and whole datasets for this experiments is available at: https://zenodo.org/records/15002575. If you are interested only in the dataset you can find it here: https://zenodo.org/records/15002371.

自然语言推理西班牙语因果关系多文体数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。