arXiv:2603.04319cs.CL2026-03ACL被引 2

用图检索+反思式提示,让大模型更准地推理事件因果。

AILS-NTUA at SemEval-2026 Task 12: Graph-Based Retrieval and Reflective Prompting for Abductive Event Reasoning

  • 三阶段流程:图检索找线索,大模型推理,后处理校正一致性。
  • 准确率0.95,击败14个模型,跨模型错误模式高度一致。
  • 揭示因果推理共性缺陷,适合研究大模型推理偏差的学者。

我们提出一种在SemEval-2026任务12中夺冠的三阶段系统,用于归纳性事件推理:结合基于图的检索、通过反思式提示优化的大模型推理,以及事后一致性强化。该系统在评估阶段排行榜上排名第一,准确率达0.95。对14个模型(7个模型家族)的跨模型错误分析揭示了三种共有的归纳偏差:因果链不完整、偏好近因、显著性偏差;这些偏差在多标签因果推理中的跨家族收敛(原因数量减少51%),表明存在系统性而非模型特异性的失败模式。

原文摘要 · Abstract (English)

We present a winning three-stage system for SemEval 2026 Task~12: Abductive Event Reasoning that combines graph-based retrieval, LLM-driven abductive reasoning with prompt design optimized through reflective prompt evolution, and post-hoc consistency enforcement; our system ranks first on the evaluation-phase leaderboard with an accuracy score of 0.95. Cross-model error analysis across 14 models (7~families) reveals three shared inductive biases: causal chain incompleteness, proximate cause preference, and salience bias, whose cross-family convergence (51\% cause-count reduction) indicates systematic rather than model-specific failure modes in multi-label causal reasoning.

因果推理大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。