arXiv:2412.10271cs.CL2024-12Transactions of th…被引 44

评估大模型生成语言的多样性,发现其远未达到人类水平。

Benchmarking Linguistic Diversity of Large Language Models

  • 从词汇、句法、语义三维度构建语言多样性评估框架。
  • 多款主流大模型在各维度上均显著低于人类语言多样性。
  • 训练数据与部署策略会显著影响生成语言的丰富程度。

大型语言模型(LLMs)的发展与评估长期聚焦于任务解决能力,近期模型甚至在某些领域超越人类表现。然而,这种关注常忽视机器生成语言是否具备人类层面的多样性,包括词汇选择、句法构造和意义表达,这引发了对语言生成基本问题是否已充分解决的质疑。本文强调评估语言模型对人类语言丰富性的保持至关重要,鉴于由大模型生成或辅助生成的在线内容呈显著上升趋势。我们提出一个涵盖词汇、句法和语义维度的综合性评估框架,对多个前沿大模型进行了全方位的多样性基准测试,并深入开展了句法多样性的案例研究。最后,分析了不同训练与部署选择如何影响大模型输出的语言多样性。

原文摘要 · Abstract (English)

The development and evaluation of Large Language Models (LLMs) has primarily focused on their task-solving capabilities, with recent models even surpassing human performance in some areas. However, this focus often neglects whether machine-generated language matches the human level of diversity, in terms of vocabulary choice, syntactic construction, and expression of meaning, raising questions about whether the fundamentals of language generation have been fully addressed. This paper emphasizes the importance of examining the preservation of human linguistic richness by language models, given the concerning surge in online content produced or aided by LLMs. We propose a comprehensive framework for evaluating LLMs from various linguistic diversity perspectives including lexical, syntactic, and semantic dimensions. Using this framework, we benchmark several state-of-the-art LLMs across all diversity dimensions, and conduct an in-depth case study for syntactic diversity. Finally, we analyze how different development and deployment choices impact the linguistic diversity of LLM outputs.

语言多样性大模型评估生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。