arXiv:2605.26104cs.CV2026-05被引 1

让多模态大模型通过实体视觉证据实现跨域视频时间定位,提升泛化能力。

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

论文配图:EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
图 1 · 摘自论文原文
  • 用实体锚定机制引导模型关注视频中的具体对象,而非依赖数据集捷径。
  • 在跨域测试中显著提升鲁棒性,同时保持原有领域性能不下降。
  • 适合需要强泛化能力的视频理解任务,如跨场景智能监控、多源视频检索。

为视频时间定位(VTG)微调多模态大模型(MLLM)常能提升本域表现,但在领域迁移时性能急剧下降。我们发现,问题根源不仅在于未见查询概念,更在于视觉领域差异,导致模型无法将已学的时间定位知识与内在的实体注意力能力关联。为此,提出EVIDENT——一种参数高效的适应框架,通过显式视觉实体证据路由VTG适配过程,将时间定位锚定在预训练MLLM的固有实体注意力之上。该框架包含三部分:(i) 实体瓶颈适配器,将密集视觉标记压缩为紧凑的实体级槽位;(ii) 实体绑定蒸馏损失,向语义无结构的MLLM视觉空间注入物体先验,引导每个槽位绑定到连贯实体;(iii) 实体到证据门控机制,利用捕获的实体作为证据,引导模型定位包含查询相关实体的片段。三者协同使微调依赖于实体锚定证据,而非脆弱的数据集捷径。跨域基准实验表明,EVIDENT在保持竞争力的本域性能的同时,显著提升跨域鲁棒性,仅需少量参数开销。结果表明,实体级定位是可泛化时间定位的有效归纳偏置。

原文摘要 · Abstract (English)

Fine-tuning MLLMs for Video Temporal Grounding (VTG) often improves in-domain performance but degrades sharply under domain shift. In this work, we find that this failure is primarily driven not just by unseen query concepts, but by visual domain shift, which prevents the model from coupling its learned temporal localization knowledge with its inherent entity-attention capability. To address this, we introduce EVIDENT, a parameter-efficient adaptation framework that anchors temporal grounding in the inherent entity-attention of pre-trained MLLMs by routing VTG adaptation through explicit visual entity evidence. EVIDENT consists of three components: (i) an Entity Bottleneck Adapter that transforms dense visual tokens into compact entity-level slots, (ii) an Entity-Binding Distillation loss that instills objectness priors into the semantically unstructured MLLM visual space, guiding each slot to bind to a coherent entity, and (iii) an Entity-to-eVidence gating mechanism that leverages the captured entities as evidence, steering the model to localize moments containing query-relevant entities. Together, these components enable VTG fine-tuning to rely on entity-grounded evidence rather than brittle dataset shortcuts. Experiments on cross-domain VTG benchmarks show that EVIDENT consistently improves out-of-domain robustness while preserving competitive in-domain performance with modest parameter overhead. These results suggest that entity-level grounding is an effective inductive bias for generalizable temporal localization.

视频定位多模态泛化能力实体对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。