arXiv:2409.14324cs.CLcs.AI2024-09EMNLP被引 7

用电影套路测试大模型叙事推理能力,发现其表现不佳且易受干扰。

Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses

  • 通过电影剧情中的套路设计评估模型抽象推理能力
  • 新方法使F1分数提升11.8点,显著改善推理效果
  • 揭示思维链提示会引发幻觉,尤其对隐含套路更敏感

具备思维链提示(CoT)的大语言模型在数学、常识和逻辑等事实性内容中展现出强大的多步推理能力,但在需要更高抽象能力的叙事推理领域表现尚不明确。本研究利用电影简介中的套路(trope)来评估前沿LLMs的抽象推理能力,发现其性能普遍偏低。为此提出一种基于套路的查询方法,将F1分数提升11.8个百分点。此外,尽管以往研究认为CoT有助于多步推理,本研究却发现,在叙事内容中,CoT反而可能引发幻觉,导致GPT-4性能下降。我们还引入对抗性注入方法,在无显式套路的电影简介中嵌入相关文本标记,揭示了CoT对这类注入的高度敏感性。全面分析为未来研究提供了重要方向。

原文摘要 · Abstract (English)

Large language models (LLMs) equipped with chain-of-thoughts (CoT) prompting have shown significant multi-step reasoning capabilities in factual content like mathematics, commonsense, and logic. However, their performance in narrative reasoning, which demands greater abstraction capabilities, remains unexplored. This study utilizes tropes in movie synopses to assess the abstract reasoning abilities of state-of-the-art LLMs and uncovers their low performance. We introduce a trope-wise querying approach to address these challenges and boost the F1 score by 11.8 points. Moreover, while prior studies suggest that CoT enhances multi-step reasoning, this study shows CoT can cause hallucinations in narrative content, reducing GPT-4's performance. We also introduce an Adversarial Injection method to embed trope-related text tokens into movie synopses without explicit tropes, revealing CoT's heightened sensitivity to such injections. Our comprehensive analysis provides insights for future research directions.

叙事推理大模型评测思维链对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。