首个针对视频大模型的反事实鲁棒性评测基准,揭示模型在篡改视频下的脆弱性。
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
- 构建动态分布外反事实视频测试集,通过风格、物体、背景等编辑生成多样化样本。
- 8个主流模型在该基准上性能显著下降,平均退化严重,验证了现有模型鲁棒性不足。
- 用反事实数据微调可提升21.73%表现,适合关注视频理解鲁棒性的研究者使用。
近期,多模态大语言模型(MLLMs)在各类视频理解任务中展现出显著性能。然而,其在面对被篡改视频内容时的鲁棒性仍缺乏系统评估。本文提出Ro-Bench,首个针对动态分布外(OOD)反事实视频测试集的MLLM评估基准。Ro-Bench通过编辑风格、物体、背景及其组合,引入高质量、多样且时间相关的视频数据。我们评估了8个近期视频MLLMs,发现模型在反事实视频上性能大幅下降。进一步表明,使用反事实数据微调可显著提升鲁棒性:在Ro-Bench上性能提升21.73%,在MVBench的20个任务上平均提升12.78%。结果证实反事实数据对增强视频理解能力的有效性。代码与数据将很快公开。
原文摘要 · Abstract (English)
Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, remains largely unexplored. In this paper, we introduce Ro-Bench, the first benchmark for evaluating MLLMs on dynamic out-of-distribution (OOD) counterfactual video test sets. Ro-Bench incorporates high-quality, diverse and temporally relevant video data, by editing Style, Object, Background and their compositions. We evaluated eight recent video MLLMs and found that current models exhibit substantial performance degradation on Ro-Bench when exposed to counterfactual video content. Furthermore, we demonstrate that fine-tuning MLLMs with counterfactual data enhances robustness, achieving a 21.73% performance increase on Ro-Bench and a 12.78% improvement across 20 tasks in the MVBench dataset. These findings underscore the effectiveness of counterfactual data in enhancing the video understanding ability of MLLMs. The code and data will be released shortly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。