arXiv:2605.27894cs.CV2026-05AAAI被引 19

提出统一模型处理视频与语言输入不完整问题,提升实际应用中的鲁棒性。

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

论文配图:Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
图 1 · 摘自论文原文
  • 设计可处理缺失模态的统一视觉语言模型架构
  • 在多任务上显著提升不完整输入下的性能表现
  • 适合需应对传感器失效的真实场景应用

视频-语言模型(VLMs)在多种计算机视觉任务中展现出强大的多模态推理能力。然而,现有VLMs通常针对特定任务且假设视频与语言输入均完整。现实中,由于传感器失效(如摄像头因隐私问题关闭),常出现模态缺失数据,导致训练与测试分布不一致。尽管不完整输入可能影响模型泛化能力甚至引发训练失败,其对VLM安全性与可信度的潜在风险尚未被充分关注。为此,本文首次提出一种统一的不完整视频-语言模型,用于处理多模态输入缺失问题。大量实验表明,该方法可作为即插即用模块,有效提升已有模型在多种多模态任务中的表现。

原文摘要 · Abstract (English)

Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due to deactivated sensors (e.g., cameras are unavailable due to data privacy), yielding modality-incomplete data and leading to inconsistency between training and testing data. While straightforward incomplete input can boast training generalization-ability and lead to training failure, its potential risks to VLMs regarding safety and trustworthiness have been largely neglected. To this end, we make the first attempt to propose a unified incomplete video-language model to process the incomplete multi-modal inputs. Extensive experimental results show that our method can serve as a plug-and-play module for previous works to improve their performance in various multi-modal tasks.

多模态视频生成鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。