探究大模型对词语与句子语义的理解能力,揭示其认知机制。
On the Semantics of Large Language Models
- 结合弗雷格与罗素的语义理论,分析大模型内部语言表征
- 发现大模型在词句层面具备一定语义理解能力
- 适合对大模型认知本质感兴趣的科研人员
大型语言模型(如ChatGPT)通过技术手段展现了复制人类语言能力的潜力,涵盖文本生成到对话交互。然而,这些系统在多大程度上真正理解语言仍存在争议。本文聚焦于大模型在词与句层面的语义问题,通过考察其内部工作机制及生成的语言表征,并借鉴弗雷格与罗素的经典语义理论,获得对大模型潜在语义能力更细致的认知图景。
原文摘要 · Abstract (English)
Large Language Models (LLMs) such as ChatGPT demonstrated the potential to replicate human language abilities through technology, ranging from text generation to engaging in conversations. However, it remains controversial to what extent these systems truly understand language. We examine this issue by narrowing the question down to the semantics of LLMs at the word and sentence level. By examining the inner workings of LLMs and their generated representation of language and by drawing on classical semantic theories by Frege and Russell, we get a more nuanced picture of the potential semantic capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。