首个专为文学细读设计的评测基准,测试大模型的解读推理能力
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning
- 构建1331道细读题,分三阶段模拟从风格提取到多跳推理的过程
- 顶尖大模型在细读任务上准确率49.7%至69.7%,低于人类专家
- 适合评估大模型在文学分析中的理解深度,尤其关注推理与上下文关联
每年有数百万篇大学英语课程论文被撰写与评分。学生需通过细读——即从文本中提取细节以支持论点——来分析文学与文化作品。尽管细读被视为批判性思维基础并广泛用于高校课程,但其尚未在大语言模型(LLMs)上得到系统评估,且跨学科基准如MMLU也未涵盖文学领域。为此,我们提出KRISTEVA,首个用于评估解释性推理的细读基准,包含1331道基于课堂数据改编的多项选择题。通过三个逐步增加难度的任务层级,模拟细读过程的不同环节:1)提取文体特征,2)从参数化知识中检索相关背景信息,3)在文体与外部语境间进行多跳推理。基线结果显示,虽顶尖大模型具备一定大学水平的细读能力(准确率49.7%–69.7%),但在11项任务中的10项仍落后于经验丰富的真人评分者。
原文摘要 · Abstract (English)
Each year, tens of millions of essays are written and graded in college-level English courses. Students are asked to analyze literary and cultural texts through a process known as close reading, in which they gather textual details to formulate evidence-based arguments. Despite being viewed as a basis for critical thinking and widely adopted as a required element of university coursework, close reading has never been evaluated on large language models (LLMs), and multi-discipline benchmarks like MMLU do not include literature as a subject. To fill this gap, we present KRISTEVA, the first close reading benchmark for evaluating interpretive reasoning, consisting of 1331 multiple-choice questions adapted from classroom data. With KRISTEVA, we propose three progressively more difficult sets of tasks to approximate different elements of the close reading process, which we use to test how well LLMs may seem to understand and reason about literary works: 1) extracting stylistic features, 2) retrieving relevant contextual information from parametric knowledge, and 3) multi-hop reasoning between style and external contexts. Our baseline results find that, while state-of-the-art LLMs possess some college-level close reading competency (accuracy 49.7% - 69.7%), their performances still trail those of experienced human evaluators on 10 out of our 11 tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。