用故事测试大模型记忆与推理能力,发现当前评估易被记忆干扰
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
- 以'故事梗概'替代原文,避免模型依赖精确记忆
- 在热门作品上准确率从96%降至72%,任务整体下降18%
- 适合关注模型真实理解力的研究者与评测人员
近期大型语言模型在角色理解任务中表现优异,如分析虚构角色的动机、性格与关系。然而,这些模型广泛预训练于大量文本,可能依赖对流行虚构作品的机械记忆而非真正理解。本文提出,角色理解应基于'梗概记忆'(捕捉核心意义),而非'逐字记忆'(字符串完全匹配)。为此设计一种简单有效的方法,在角色理解评测中抑制机械记忆,同时保留隐含理解线索。实验显示,该方法将热门作品上的准确率从96%降至72%,各类角色理解任务最高下降18%。结果表明,现有基准存在数据污染问题,常衡量记忆而非真实理解。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) have shown impressive performance in character understanding tasks, such as analyzing the roles, personalities, and relationships of fictional characters. However, the extensive pre-training corpora used by LLMs raise concerns that they may rely on memorizing popular fictional works rather than genuinely understanding and reasoning about them. In this work, we argue that 'gist memory'-capturing essential meaning - should be the primary mechanism for character understanding tasks, as opposed to 'verbatim memory' - exact match of a string. We introduce a simple yet effective method to mitigate mechanized memorization in character understanding evaluations while preserving the essential implicit cues needed for comprehension and reasoning. Our approach reduces memorization-driven performance on popular fictional works from 96% accuracy to 72% and results in up to an 18% drop in accuracy across various character understanding tasks. These findings underscore the issue of data contamination in existing benchmarks, which often measure memorization rather than true character understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。