构建真实事件因果推理数据集,评估大模型找直接原因的能力
SemEval-2026 Task 12: Abductive Event Reasoning: Towards Real-World Event Causal Inference for Large Language Models
- 设计多选题形式,让模型从证据中推断真实事件的直接原因
- 122支队伍提交518份结果,展现大模型在复杂因果推理上的不足
- 适合研究因果推理、多文档理解的学者和工程师参考
理解现实事件发生的原因对自然语言处理和实际决策都至关重要,但现有研究仍缺乏对证据丰富场景下直接因果推理的深入探索。为此,我们组织了 SemEval-2026 Task 12:归纳事件推理(Abductive Event Reasoning, AER)。该任务要求系统从支持性证据中识别目标事件的最可能直接原因。AER 被构造成一个基于证据的多选基准,涵盖真实因果推理的关键挑战,包括分散证据、间接背景因素以及语义相关但非因果的干扰项。共有 122 名参与者提交了 518 份系统结果。本文介绍了任务设计、数据构建流程、评估方案及系统表现。AER 为真实事件的归纳推理提供了聚焦基准,揭示了未来因果推理与多文档理解的研究挑战。
原文摘要 · Abstract (English)
Understanding why real-world events occur is important for both natural language processing and practical decision-making, yet direct-cause inference remains underexplored in evidence-rich settings. To address this gap, we organized SemEval-2026 Task 12: Abductive Event Reasoning (AER).\footnote{The task data is available at https://github.com/sooo66/semeval2026-task12-dataset.git} The task asks systems to identify the most plausible direct cause of a target event from supporting evidence. We formulate AER as an evidence-grounded multiple-choice benchmark that captures key challenges of real-world causal reasoning, including distributed evidence, indirect background factors, and semantically related but non-causal distractors. The shared task attracted 122 participants and received 518 submissions. This paper presents the task formulation, dataset construction pipeline, evaluation setup, and system results. AER provides a focused benchmark for abductive reasoning over real-world events and highlights challenges for future work on causal reasoning and multi-document understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。