arXiv:2503.06137cs.CL2025-03被引 3

评测大模型在话语连贯性上的表现,发现不同模型差异明显。

Evaluating Discourse Cohesion in Pre-trained Language Models

  • 构建多类句间连贯性测试集,覆盖相邻与非相邻句子关系。
  • 对比多个预训练模型,发现其连贯性能力存在显著差异。
  • 为未来研究提供评估基准,适合关注语言生成质量的读者。

大型预训练神经模型在自然语言处理中取得了显著成功,激发了从多个角度分析其能力的研究热潮。本文提出一个测试套件,用于评估预训练语言模型的连贯性能力。该套件包含相邻与非相邻句子之间的多种连贯现象。我们尝试在这些现象上对比不同预训练语言模型的表现,并分析实验结果,期望未来能更多关注话语连贯性问题。

原文摘要 · Abstract (English)

Large pre-trained neural models have achieved remarkable success in natural language process (NLP), inspiring a growing body of research analyzing their ability from different aspects. In this paper, we propose a test suite to evaluate the cohesive ability of pre-trained language models. The test suite contains multiple cohesion phenomena between adjacent and non-adjacent sentences. We try to compare different pre-trained language models on these phenomena and analyze the experimental results,hoping more attention can be given to discourse cohesion in the future.

语言模型连贯性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。