arXiv:2603.06683cs.CV2026-03

ECHO通过多智能体协作显式构建事件超图,提升多媒体事件抽取的可解释性与准确性。

ECHO: Event-Centric Hypergraph Operations via Multi-Agent Collaboration for Multimedia Event Extraction

  • 将事件抽取重构为共享超图上的迭代更新,中间结构清晰可查。
  • 在事件提及和论元角色上分别提升7.3和15.5 F1点,超越现有方法。
  • 适合关注可解释性、需要修正中间预测的研究者与开发者。

多媒体事件抽取(M2E2)旨在预测触发词,在文本与图像中定位论元,并将其组合成符合模式的事件记录。现有基于大模型的方法虽具潜力,但中间事件假设常隐含不清,且事件-论元链接与角色绑定紧密耦合,难以检查或修正,导致预测对早期错误敏感。为此,我们提出ECHO,一个将M2E2重构为显式多媒体事件超图(MEHG)迭代优化的多智能体框架。不同于依赖隐式线性生成,ECHO在共享超图上执行可审计的原子更新,使中间事件结构显式且可修改。此外,引入“先链接后绑定”策略,解耦事件-论元链接与角色绑定,降低结构化预测中的过早语义承诺。在M2E2基准上的大量实验表明,ECHO持续优于先前最优方法,在事件提及和论元角色上分别取得7.3和15.5 F1点的提升。

原文摘要 · Abstract (English)

Multimedia event extraction (M2E2) aims to predict triggers, ground arguments across text and images, and then assemble them into schema-consistent event records. Recent LLM-based approaches have shown strong potential for M2E2, but their intermediate event hypotheses often remain implicit, and event-argument linking is still tightly coupled with role binding. This leaves little opportunity to inspect or revise intermediate event hypotheses and makes predictions brittle to early errors. To bridge this gap, we present ECHO, a multi-agent framework that reframes M2E2 as iterative refinement over an explicit Multimedia Event Hypergraph (MEHG). Instead of relying on implicit linear generation, ECHO performs auditable atomic updates over a shared hypergraph, making intermediate event structures explicit and revisable. Furthermore, we introduce a Link-then-Bind strategy that decouples event-argument linking from role binding, reducing premature semantic commitment during structured prediction. Extensive experiments on the M2E2 benchmark show that ECHO consistently outperforms prior state-of-the-art approaches, achieving gains of 7.3 and 15.5 F1 points on event mention and argument role, respectively.

事件抽取多智能体超图可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。