构建叙事理解评估框架,揭示现有评测短板
NarraBench: A Comprehensive Framework for Narrative Benchmarking
- 提出理论驱动的叙事任务分类体系
- 78个基准中仅27%任务被有效覆盖
- 聚焦主观视角、风格等无标准答案的评估
我们提出NarraBench,一个基于理论的叙事理解任务分类体系,以及对该领域78个现有基准的系统性调查。研究发现,当前评估存在显著空白:仅27%的叙事任务被现有基准充分覆盖,而叙事事件、风格、视角和揭示等关键方面几乎未被纳入评估。此外,亟需开发能衡量叙事中固有主观性与视角性的评测工具,因为这些方面通常不存在唯一正确答案。本研究的分类体系、调查结果与方法论对希望测试大模型叙事理解能力的自然语言处理研究者具有重要参考价值。
原文摘要 · Abstract (English)
We present NarraBench, a theory-informed taxonomy of narrative-understanding tasks, as well as an associated survey of 78 existing benchmarks in the area. We find significant need for new evaluations covering aspects of narrative understanding that are either overlooked in current work or are poorly aligned with existing metrics. Specifically, we estimate that only 27% of narrative tasks are well captured by existing benchmarks, and we note that some areas -- including narrative events, style, perspective, and revelation -- are nearly absent from current evaluations. We also note the need for increased development of benchmarks capable of assessing constitutively subjective and perspectival aspects of narrative, that is, aspects for which there is generally no single correct answer. Our taxonomy, survey, and methodology are of value to NLP researchers seeking to test LLM narrative understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。