构建长文本指代消解新基准,揭示大模型在复杂指代中的理解短板。
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
- 采用多选题形式,结合长叙事与去噪策略,提升指代任务难度。
- 封闭模型最高准确率81.9%(GPT-4o),开放模型落后20-30%。
- 嵌套结构易混淆实体,代词指代比名词指代更难解析。
当前大模型在指代消解任务上的评估发现,传统输出格式与评价指标未能充分反映其指代理解能力。为此,我们提出IdentifyMe,一个以多选题形式呈现的指代消解基准,适用于评估大模型。该基准包含长篇叙述,并通过启发式方法排除易于识别的指代项,使任务更具挑战性。其数据集融合多种指代类型与对应实体,支持对模型性能的细粒度分析。我们在IdentifyMe上评估了闭源与开源大模型,发现顶尖闭源模型(如GPT-4o)表现显著优于开源模型(100亿参数以下),差距达20%-30%。研究还表明,表面信息有限的代词指代远比名词指代更难处理;当多个指代在嵌套结构中重叠时,模型常发生实体混淆。最高得分模型(GPT-4o)达到81.9%准确率,表明当前先进模型具备较强指代能力,但仍存在改进空间。
原文摘要 · Abstract (English)
Recent evaluations of LLMs on coreference resolution have revealed that traditional output formats and evaluation metrics do not fully capture the models' referential understanding. To address this, we introduce IdentifyMe, a new benchmark for mention resolution presented in a multiple-choice question (MCQ) format, commonly used for evaluating LLMs. IdentifyMe features long narratives and employs heuristics to exclude easily identifiable mentions, creating a more challenging task. The benchmark also consists of a curated mixture of different mention types and corresponding entities, allowing for a fine-grained analysis of model performance. We evaluate both closed- and open source LLMs on IdentifyMe and observe a significant performance gap (20-30%) between the state-of-the-art sub-10B open models vs. closed ones. We observe that pronominal mentions, which have limited surface information, are typically much harder for models to resolve than nominal mentions. Additionally, we find that LLMs often confuse entities when their mentions overlap in nested structures. The highest-scoring model, GPT-4o, achieves 81.9% accuracy, highlighting the strong referential capabilities of state-of-the-art LLMs while also indicating room for further improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。