arXiv:2505.14405cs.CV2025-05ACL被引 2

提出新评测基准,提升大模型对视频时序不一致的鲁棒性

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

  • 设计视觉与文本双模态时序扰动评测集
  • 16个主流模型在对抗环境下依赖文本而非时序动态
  • 提出全景直接偏好优化,同步融合多模态特征

大型多模态模型(LMMs)在通用视频理解基准上表现优异,但其时序分析能力的鲁棒性尚未被充分考察。为此,我们提出了一个新的时序鲁棒性评测基准(TemRobBench),分别在视觉和文本模态引入时序不一致扰动以评估模型表现。我们评估了16个主流LMMs,发现它们在对抗环境中过度依赖先验知识和文本上下文,忽视视频中的实际时序动态。为缓解此问题,我们设计了全景直接偏好优化(PanoDPO),促使模型同时整合视觉与语言特征偏好。实验表明,PanoDPO能有效提升模型在时序分析中的鲁棒性与可靠性。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, we propose a novel temporal robustness benchmark (TemRobBench), which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. We evaluate 16 mainstream LMMs and find that they exhibit over-reliance on prior knowledge and textual context in adversarial environments, while ignoring the actual temporal dynamics in the video. To mitigate this issue, we design panoramic direct preference optimization (PanoDPO), which encourages LMMs to incorporate both visual and linguistic feature preferences simultaneously. Experimental results show that PanoDPO can effectively enhance the model's robustness and reliability in temporal analysis.

多模态视频理解鲁棒性偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。