揭示多模态事件抽取评估中的三大陷阱,提醒研究者警惕结果高估风险。
Evaluation Pitfalls and Challenges in Multimedia Event Extraction

- 首次系统分析多模态事件抽取的评估漏洞
- 微小评估选择可导致性能差异显著夸大
- 适合关注评估公平性与可比性的研究者
多模态事件抽取旨在跨文本与图像等多模态数据联合识别事件及其论元,以支持更全面的事件理解。尽管近期研究宣称持续且显著进步,但这些结果的可靠性与可比性高度依赖于一致且严格的评估标准。本文首次对多模态事件抽取中的评估陷阱进行系统分析,识别出三大问题来源:数据处理不一致、任务假设不统一、评估设置过于宽松。通过在严格评估框架下的系列控制实验,我们证明微小的评估选择即可引发显著性能波动,导致模型跨模态真实事件定位能力被严重高估。研究强调亟需建立可比的评估标准,并推动该领域向更严谨的评估范式转变。
原文摘要 · Abstract (English)
Multimedia event extraction aims to jointly identify events and their arguments across multiple modalities, such as text and images, to support more comprehensive event understanding. While recent work reports steady and substantial progress, the reliability and comparability of these results critically depend on consistent and rigorous evaluation. In this work, we present the first systematic analysis of evaluation pitfalls in multimedia event extraction and identify three major sources of issues: inconsistent data processing, inconsistent task assumptions, and overly relaxed evaluation settings. We demonstrate, through a series of controlled experiments under a strict evaluation framework, that minor evaluation choices can cause large performance variations and lead to overestimation of a model's ability to ground real-world events across modalities. Our findings highlight the need for comparable evaluation standards and encourage a shift toward more rigorous evaluation in multimedia event extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。