arXiv:2604.23173cs.CV2026-04中稿 · CVPR

通过多模态实体消歧,提升视频事件理解中角色识别的一致性。

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

论文配图:One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
图 1 · 摘自论文原文
  • 构建多阶段实体消歧框架,融合文本描述与视觉特征
  • 在CIDEr和LEA指标上提升2.5%和7%,视觉定位准确率提升18%
  • 适用于需要跨镜头角色一致识别的视频理解任务

视频情境识别(VidSitu)旨在解决视频中“谁对谁做了什么,用什么,如何,何处”的复杂问题,要求对多个事件中的关键动作及关联角色进行识别。该任务需要对跨镜头、外观变化的实体进行时空定位。本文提出多模态实体消歧(MEC),通过统一文本中的实体描述与视频中的视觉定位,实现更连贯的理解。为此,我们设计了分阶段的CineMEC方法,在无显式定位标注的情况下,将事件角色提及组与视觉实体聚类相匹配。该方法利用视觉定位与字幕生成之间的协同作用,两者互促提升。我们扩展了VidSitu数据集以包含定位标注。实验表明,相比以往仅关注描述生成的工作,CineMEC在字幕生成(+2.5% CIDEr,+7% LEA)和视觉定位(+18% HOTA)上均取得显著提升。

原文摘要 · Abstract (English)

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient actions and associated short descriptions for event roles across multiple events. Grounding with VidSitu requires spatio-temporal localization of key entities across shots and varied appearances. We posit that coherent video understanding requires consistent identification of entities that play different roles. We propose Multimodal Entity Coreference (MEC) to unite entity descriptions in text with grounding across the video. Towards this, we introduce CineMEC, a multi-stage approach that unites event role mention groups with visual clusters of entities, without explicit grounding supervision during training. Our approach is designed to exploit the synergy between visual grounding and captioning, where improving one influences the other and vice versa. For evaluation, we extend the VidSitu dataset with grounding annotations. While previous work focuses primarily on descriptions, CineMEC improves consistency across both: captioning (+2.5% CIDEr, +7% LEA) and visual grounding (+18% HOTA).

视频理解实体消歧多模态场景识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。