arXiv:2503.18637cs.CV2025-03CVPR被引 6

用文本描述消除视频数据集偏见,提升模型评估可靠性。

Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks

  • 通过多模态模型生成视频帧级文本描述,识别并去除对象、时间等偏差。
  • 分析12个主流视频数据集,构建去对象偏见的测试集,发现多数模型依赖表面特征。
  • 开源描述数据与去偏测试集,助力更公平的视频理解研究。

我们提出一个新的「通过文本描述去偏(UTD)」视频基准,基于现有视频分类与检索数据集中的无偏子集,以实现对视频理解能力更可靠的评估。当前视频基准可能受多种表示偏见影响,例如物体偏见或单帧偏见,仅识别物体或仅使用单帧即可正确预测。我们利用视觉语言模型(VLMs)和大语言模型(LLMs)分析并消除这些偏见。具体地,生成视频每帧的文本描述,筛选特定信息(如仅物体),并从三个维度检验表示偏见:1)概念偏见——特定概念(如物体)是否足以完成预测;2)时间偏见——时间信息是否对预测有贡献;3)常识推理与数据集偏见——预测是依赖零样本推理还是数据集关联。我们系统分析了12个流行的视频分类与检索数据集,并为这些数据集创建新的去物体偏见的测试集。此外,我们在原始和去偏测试集上对30个先进视频模型进行基准测试,并分析模型中的偏见。为促进未来更鲁棒的视频理解基准与模型的发展,我们发布了:「UTD-descriptions」——包含各数据集丰富结构化描述的数据集,以及「UTD-splits」——去物体偏见的测试集。

原文摘要 · Abstract (English)

We propose a new "Unbiased through Textual Description (UTD)" video benchmark based on unbiased subsets of existing video classification and retrieval datasets to enable a more robust assessment of video understanding capabilities. Namely, we tackle the problem that current video benchmarks may suffer from different representation biases, e.g., object bias or single-frame bias, where mere recognition of objects or utilization of only a single frame is sufficient for correct prediction. We leverage VLMs and LLMs to analyze and debias benchmarks from such representation biases. Specifically, we generate frame-wise textual descriptions of videos, filter them for specific information (e.g. only objects) and leverage them to examine representation biases across three dimensions: 1) concept bias - determining if a specific concept (e.g., objects) alone suffice for prediction; 2) temporal bias - assessing if temporal information contributes to prediction; and 3) common sense vs. dataset bias - evaluating whether zero-shot reasoning or dataset correlations contribute to prediction. We conduct a systematic analysis of 12 popular video classification and retrieval datasets and create new object-debiased test splits for these datasets. Moreover, we benchmark 30 state-of-the-art video models on original and debiased splits and analyze biases in the models. To facilitate the future development of more robust video understanding benchmarks and models, we release: "UTD-descriptions", a dataset with our rich structured descriptions for each dataset, and "UTD-splits", a dataset of object-debiased test splits.

视频理解去偏数据集构建多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。