arXiv:2411.01539cs.CLcs.LG2024-11被引 3

分析大模型错误答案的规律,发现不同模型错误高度相关。

LLMs and the Madness of Crowds

  • 通过错误模式分析模型间相似性
  • 错误分布非随机,存在系统性关联
  • 构建错误相关性的分类体系

我们研究了大语言模型(LLMs)在评估过程中产生的错误答案模式。这些错误表现出每个模型特有的高度非直观行为。通过分析这些模式,我们测量了不同模型间的相似性,并基于错误相关性构建了一个分类体系。研究发现,错误回答并非随机分布,而是在模型间系统性地相关,为理解大模型内部结构及相互关系提供了新视角。

原文摘要 · Abstract (English)

We investigate the patterns of incorrect answers produced by large language models (LLMs) during evaluation. These errors exhibit highly non-intuitive behaviors unique to each model. By analyzing these patterns, we measure the similarities between LLMs and construct a taxonomy that categorizes them based on their error correlations. Our findings reveal that the incorrect responses are not randomly distributed but systematically correlated across models, providing new insights into the underlying structures and relationships among LLMs.

大模型错误分析模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。