arXiv:2509.14571cs.HCcs.AI2025-09中稿 · IEEE VIS 2025被引 1

用可视化分析工具评估视觉语言模型在数据损坏下的鲁棒性

VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models

  • 构建可视化框架,分析不同数据损坏对模型的影响
  • 识别性能下降样本,指导数据增强策略设计
  • 适合模型开发者与可信人工智能研究者使用

视觉语言(VL)模型在多模态信息理解方面展现出巨大潜力,但在分布偏移下性能常显著下降,亟需评估和提升其对现实数据损坏的鲁棒性。尽管基准数据集和数据增强(DA)已有进展,但对模型行为的深层理解不足,且需专家经验与反复试验来探索数据模式。鉴于可视化在解释复杂模型和大规模数据分析中的成功,我们提出VisMoDAl——一个用于评估VL模型在各类数据损坏下的鲁棒性,并识别表现不佳样本以指导有效数据增强策略的可视化分析框架。基于文献综述与专家讨论,该框架支持从特定损坏类型到任务驱动的模型行为与数据子集分析。相比传统方法,它使用户能推理损坏对模型的影响,促进模型行为理解与数据增强策略制定。通过图像字幕任务的案例研究与定量评估验证了系统有效性。

原文摘要 · Abstract (English)

Vision-language (VL) models have shown transformative potential across various critical domains due to their capability to comprehend multi-modal information. However, their performance frequently degrades under distribution shifts, making it crucial to assess and improve robustness against real-world data corruption encountered in practical applications. While advancements in VL benchmark datasets and data augmentation (DA) have contributed to robustness evaluation and improvement, there remain challenges due to a lack of in-depth comprehension of model behavior as well as the need for expertise and iterative efforts to explore data patterns. Given the achievement of visualization in explaining complex models and exploring large-scale data, understanding the impact of various data corruption on VL models aligns naturally with a visual analytics approach. To address these challenges, we introduce VisMoDAl, a visual analytics framework designed to evaluate VL model robustness against various corruption types and identify underperformed samples to guide the development of effective DA strategies. Grounded in the literature review and expert discussions, VisMoDAl supports multi-level analysis, ranging from examining performance under specific corruptions to task-driven inspection of model behavior and corresponding data slice. Unlike conventional works, VisMoDAl enables users to reason about the effects of corruption on VL models, facilitating both model behavior understanding and DA strategy formulation. The utility of our system is demonstrated through case studies and quantitative evaluations focused on corruption robustness in the image captioning task.

视觉语言模型鲁棒性评估可视化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。