首个全面评估视频大模型可信度的基准,揭示其在真实场景中的多重风险。
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
- 构建5维综合评测框架,覆盖真实性、鲁棒性、安全、公平与隐私
- 23个主流模型在动态场景理解与跨模态干扰下表现普遍不足
- 开源模型偶有超越,但闭源模型整体更可信,需加强数据多样性
面向视频理解的多模态大模型(videoLLMs)虽在处理复杂时空数据方面取得进展,但事实错误、有害内容、偏见、幻觉和隐私风险严重削弱其可靠性。本研究提出Trust-videoLLMs,首个综合性基准,评估23个先进videoLLMs(5个商用,18个开源),涵盖真实性、鲁棒性、安全、公平与隐私五个维度。该框架包含30项任务,使用改编、合成及标注视频,评估时空风险、时间一致性与跨模态影响。结果表明,模型在动态场景理解、跨模态扰动鲁棒性及现实风险缓解方面存在显著缺陷。尽管部分开源模型表现优异,但总体上闭源模型更具可信度,且规模扩大并不保证性能提升。研究强调需增强训练数据多样性与多模态对齐能力。Trust-videoLLMs为标准化可信度评估提供公开可扩展工具,填补了仅关注准确率的评测与实际应用中对鲁棒性、安全、公平与隐私需求之间的空白。
原文摘要 · Abstract (English)
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。