用统一标签集测试大模型跨语言、跨框架的语篇理解能力
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
- 构建统一语篇关系标签集,支持多语言多框架分析
- 23个大模型实验显示,多语言训练模型具备强跨语言泛化能力
- 中间层最擅长语篇层面的语言泛化,适合研究跨语言语篇模型
语篇理解对众多自然语言处理任务至关重要,但现有研究多受限于依赖特定框架的语篇表示。本文探究大语言模型(LLMs)是否能捕捉跨语言与跨框架的通用语篇知识。从两个维度展开:(1) 构建统一语篇关系标签集,促进跨语言与跨框架分析;(2) 通过探测方法评估LLMs是否编码可泛化的语篇抽象。以多语言语篇关系分类为测试任务,考察了23个不同规模与多语言能力的LLMs。结果表明,尤其是经过多语言训练的模型,能够有效在语言和框架间泛化语篇信息。分层分析显示,语篇层面的语言泛化在中间层最为显著。最后的错误分析揭示了部分关系类别的识别难点。
原文摘要 · Abstract (English)
Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations. This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks. We address this question along two dimensions: (1) developing a unified discourse relation label set to facilitate cross-lingual and cross-framework discourse analysis, and (2) probing LLMs to assess whether they encode generalizable discourse abstractions. Using multilingual discourse relation classification as a testbed, we examine a comprehensive set of 23 LLMs of varying sizes and multilingual capabilities. Our results show that LLMs, especially those with multilingual training corpora, can generalize discourse information across languages and frameworks. Further layer-wise analyses reveal that language generalization at the discourse level is most salient in the intermediate layers. Lastly, our error analysis provides an account of challenging relation classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。