arXiv:2409.06622cs.CL2024-09被引 8

探究预训练模型对意大利语句法语义信息的编码能力

Exploring Italian sentence embeddings properties through multi-tasking

  • 用大规模合成数据构建多任务框架研究嵌入表示
  • 不同任务线索在嵌入中以不同方式编码,表现不一致
  • 发现预训练嵌入未有效捕捉抽象语言概念

我们通过多任务设置,研究现有大语言模型在意大利语中编码抽象语言信息的程度。利用大规模精心构建的合成数据——多个 Blackbird Language Matrices (BLMs) 意大利语问题——来分析基于预训练语言模型构建的句子表示如何编码特定句法与语义信息。采用两级架构,分别建模句子嵌入压缩为任务相关表示,以及 BLM 任务。进一步探究能否获得能编码多个 BLM 任务所需句法与语义信息的压缩表示。尽管预期句子结构(短语/成分序列)和成分属性可在任务间共享,但性能与错误分析显示,不同任务的线索在句子嵌入中以不同方式编码,表明诸如成分或主题角色等抽象语言概念似乎并未存在于预训练句子嵌入中。

原文摘要 · Abstract (English)

We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several Blackbird Language Matrices (BLMs) problems in Italian -- and use them to study how sentence representations built using pre-trained language models encode specific syntactic and semantic information. We use a two-level architecture to model separately a compression of the sentence embeddings into a representation that contains relevant information for a task, and a BLM task. We then investigate whether we can obtain compressed sentence representations that encode syntactic and semantic information relevant to several BLM tasks. While we expected that the sentence structure -- in terms of sequence of phrases/chunks -- and chunk properties could be shared across tasks, performance and error analysis show that the clues for the different tasks are encoded in different manners in the sentence embeddings, suggesting that abstract linguistic notions such as constituents or thematic roles does not seem to be present in the pretrained sentence embeddings.

句子嵌入多任务学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。