arXiv:2504.15604cs.CLcs.AI2025-04

对比GPT-2与LLaMA-2在心理理论任务中的下一个词预测表现。

Exploring Next Token Prediction in Theory of Mind (ToM) Tasks: Comparative Experiments with GPT-2 and LLaMA-2 AI Models

  • 构建含不同复杂度上下文的短故事数据集,评估模型推理能力。
  • LLaMA-2在低温度下表现更优,准确率更高且更稳定。
  • 随着推理层级提升,模型预测差异增大,体现认知复杂性影响。

语言模型在生成连贯文本和基于提示预测下一个词方面取得显著进展。本研究比较了OpenAI的GPT-2与Meta的Llama-2-7b-chat-hf在心理理论(ToM)任务中的下一个词预测表现。数据集基于10个短故事构建,源自Explore ToM Dataset,通过GPT-4程序化插入额外句子(infills),生成不同上下文复杂度的变体。实验在四个温度设置(0.01、0.5、1.0、2.0)下进行,评估模型在零阶(追踪状态)、一阶(理解他人心理)和二阶(递归推理)三种推理层级下的预测能力。结果表明,增加上下文会轻微降低预测准确率,因复杂性和模糊性上升。LLaMA-2整体优于GPT-2,尤其在低温下表现更佳,体现出更强的置信度与稳定性。随着推理复杂度提升,模型输出分歧加剧,尤其在一阶与二阶任务中预测波动更大。研究揭示了模型架构、温度设置与上下文复杂度对下一个词预测的影响,有助于深入理解当前语言模型的优势与局限。

原文摘要 · Abstract (English)

Language models have made significant progress in generating coherent text and predicting next tokens based on input prompts. This study compares the next-token prediction performance of two well-known models: OpenAI's GPT-2 and Meta's Llama-2-7b-chat-hf on Theory of Mind (ToM) tasks. To evaluate their capabilities, we built a dataset from 10 short stories sourced from the Explore ToM Dataset. We enhanced these stories by programmatically inserting additional sentences (infills) using GPT-4, creating variations that introduce different levels of contextual complexity. This setup enables analysis of how increasing context affects model performance. We tested both models under four temperature settings (0.01, 0.5, 1.0, 2.0) and evaluated their ability to predict the next token across three reasoning levels. Zero-order reasoning involves tracking the state, either current (ground truth) or past (memory). First-order reasoning concerns understanding another's mental state (e.g., "Does Anne know the apple is salted?"). Second-order reasoning adds recursion (e.g., "Does Anne think that Charles knows the apple is salted?"). Our results show that adding more infill sentences slightly reduces prediction accuracy, as added context increases complexity and ambiguity. Llama-2 consistently outperforms GPT-2 in prediction accuracy, especially at lower temperatures, demonstrating greater confidence in selecting the most probable token. As reasoning complexity rises, model responses diverge more. Notably, GPT-2 and Llama-2 display greater variability in predictions during first- and second-order reasoning tasks. These findings illustrate how model architecture, temperature, and contextual complexity influence next-token prediction, contributing to a better understanding of the strengths and limitations of current language models.

语言模型心理理论下一个词预测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。