arXiv:2502.07747cs.CLcs.AI2025-02被引 4

测试大模型在侦探故事中找真凶的推理能力,发现名字替换会影响判断准确率。

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories

  • 构建侦探故事数据集WhoDunIt,用真实人物名替换角色名测试推理鲁棒性。
  • GPT-4系列模型在未改动文本上表现稳定,但知名人物替换导致准确率下降。
  • 适合研究大模型叙事推理能力或评估提示工程效果的研究者使用。

我们提出一个名为WhoDunIt的新数据集,用于评估大语言模型(LLM)在叙事语境下的演绎推理能力。该数据集源自开放域侦探小说和短篇故事,要求模型在阅读并理解故事后识别出真正的作案人。为评估模型鲁棒性,我们对角色名称进行了多种层级的修改,包括原名、姓名互换以及替换为大众熟知的真实或虚构人物。同时采用不同提示风格,探究提示方式对演绎推理准确性的影响。我们对当前最先进的模型(GPT-4o、GPT-4-turbo 和 GPT-4o-mini)进行了多轮测试,通过多数投票选择结果以确保可靠性。实验表明,尽管模型在未修改文本上表现稳定,但在涉及广泛认知的人物替换时,准确率明显下降。该数据集已公开。

原文摘要 · Abstract (English)

We present a novel data set, WhoDunIt, to assess the deductive reasoning capabilities of large language models (LLM) within narrative contexts. Constructed from open domain mystery novels and short stories, the dataset challenges LLMs to identify the perpetrator after reading and comprehending the story. To evaluate model robustness, we apply a range of character-level name augmentations, including original names, name swaps, and substitutions with well-known real and/or fictional entities from popular discourse. We further use various prompting styles to investigate the influence of prompting on deductive reasoning accuracy. We conduct evaluation study with state-of-the-art models, specifically GPT-4o, GPT-4-turbo, and GPT-4o-mini, evaluated through multiple trials with majority response selection to ensure reliability. The results demonstrate that while LLMs perform reliably on unaltered texts, accuracy diminishes with certain name substitutions, particularly those with wide recognition. This dataset is publicly available here.

推理评测叙事理解大模型测试侦探故事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。