用大模型自动生成并评估不同合理性的事件,提升心理语言学研究效率。
STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

- 基于大模型构建共享事件框架,仅变换单一角色实现可控合理性设计。
- 引入全局推理与评估引导优化后,高质量事件生成率从16.7%提升至75.0%。
- 在边界难例上仍需人工干预,提示人类判断对极端情况的关键作用。
事件知识涉及谁对谁做了什么。心理语言学家使用事件合理性判断来研究此类知识如何支持语言处理。为隔离合理性影响,研究需控制事件集,其中仅一个事件槽位在合理性等级间变化,其余特征保持不变。手动构建此类集合耗时费力。为此,我们提出STRIVE,一种基于大模型的框架,用于联合生成与评估跨越合理性类别(合理 vs. 不合理)和分类难度(易 vs. 难)的受控事件集。给定一个动词,STRIVE 构建共享事件框架,然后通过改变单一槽位、固定其余项,生成每种条件下的事件。在60个动词上对六种模型的实验表明,使用基础生成提示时,GPT-5.1仅16.7%情况下生成高质量集合;增加全局推理草稿与评估器引导修正后,该比例升至75.0%。更强的推理努力也提升了评估器与人类的一致性。然而,接近合理边界的情况最困难:人类分歧最大,最优评估器在不合理-难条件下准确率仅为57%,表明仍需人类介入。总体而言,STRIVE 提供了一种可扩展的方法,显著减少人工工作量,自动化心理语言学研究中的初始事件集生成与评估。
原文摘要 · Abstract (English)
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。