arXiv:2501.05552cs.CLcs.AI2025-01被引 3

评测大模型对概念随时间演变的理解能力,揭示其历史语义把握差异。

The dynamics of meaning through time: Assessment of Large Language Models

  • 设计特定提示词,分析模型对跨时期词汇的解读表现。
  • 多模型对比显示在历史语境理解上存在显著差异。
  • 适合关注AI语言演化、数字人文的研究者参考。

理解大语言模型(LLMs)如何掌握概念的历史背景及其语义演变,对推动人工智能与语言学研究至关重要。本研究旨在评估多种LLM捕捉语义时间动态的能力,特别是它们对不同时期术语的解释。我们选取多个领域的多样化词汇,采用定制化提示词,并通过客观指标(如困惑度、词数)和主观专家评估来衡量模型响应。对比分析涵盖ChatGPT、GPT-4、Claude、Bard、Gemini和Llama等主流模型。结果表明,各模型在处理历史语境与语义变迁方面表现差异明显,凸显了其在时间语义理解上的优势与局限。这些发现为改进LLM以更好应对语言演变提供了基础,对历史文本分析、AI设计及数字人文应用具有重要意义。

原文摘要 · Abstract (English)

Understanding how large language models (LLMs) grasp the historical context of concepts and their semantic evolution is essential in advancing artificial intelligence and linguistic studies. This study aims to evaluate the capabilities of various LLMs in capturing temporal dynamics of meaning, specifically how they interpret terms across different time periods. We analyze a diverse set of terms from multiple domains, using tailored prompts and measuring responses through both objective metrics (e.g., perplexity and word count) and subjective human expert evaluations. Our comparative analysis includes prominent models like ChatGPT, GPT-4, Claude, Bard, Gemini, and Llama. Findings reveal marked differences in each model's handling of historical context and semantic shifts, highlighting both strengths and limitations in temporal semantic understanding. These insights offer a foundation for refining LLMs to better address the evolving nature of language, with implications for historical text analysis, AI design, and applications in digital humanities.

大模型评测语义演化历史语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。