arXiv:2506.03589cs.CVcs.AI2025-06中稿 · ACM MM 2025被引 5

通过场景元素引导,减轻文本-视频检索中的视觉语言偏见。

BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance

  • 用场景元素提取视频关键实体与动作,增强细粒度表示。
  • 在5个主流数据集上性能优于基线,跨分布检索表现稳定。
  • 适合关注多模态公平性、细节感知检索的研究者。

文本-视频检索(TVR)系统常受数据集中存在的视觉-语言偏见影响,导致预训练模型忽略关键细节。为此,我们提出BiMa框架,旨在同时缓解视觉与文本表征中的偏见。方法首先通过识别相关实体/物体和活动,生成刻画每个视频的场景元素。视觉去偏方面,将这些场景元素融入视频嵌入,强化细粒度与显著特征;文本去偏方面,引入解耦机制将文本特征分离为内容与偏见成分,使模型能专注有意义内容并独立处理偏见信息。在五个主要TVR基准(MSR-VTT、MSVD、LSMDC、ActivityNet和DiDeMo)上的大量实验及消融研究显示,BiMa表现优异。此外,其去偏能力在跨分布检索任务中也持续得到验证。

原文摘要 · Abstract (English)

Text-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To address this, we propose BiMa, a novel framework designed to mitigate biases in both visual and textual representations. Our approach begins by generating scene elements that characterize each video by identifying relevant entities/objects and activities. For visual debiasing, we integrate these scene elements into the video embeddings, enhancing them to emphasize fine-grained and salient details. For textual debiasing, we introduce a mechanism to disentangle text features into content and bias components, enabling the model to focus on meaningful content while separately handling biased information. Extensive experiments and ablation studies across five major TVR benchmarks (i.e., MSR-VTT, MSVD, LSMDC, ActivityNet, and DiDeMo) demonstrate the competitive performance of BiMa. Additionally, the model's bias mitigation capability is consistently validated by its strong results on out-of-distribution retrieval tasks.

文本视频检索去偏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。