发现视频检索模型偏好生成视频,揭示了内容生态中的隐性偏见。
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
- 构建1.3万条真实与生成视频的基准数据集,评估检索偏见。
- 三类检索模型均更倾向生成视频,且训练集含生成内容会加剧偏见。
- 偏见源于视觉与时间双重因素,适合关注AIGC影响的研究者阅读。
随着AI生成内容(AIGC)快速发展,高质量视频的生成日益便捷,导致网络充斥各类视频内容。然而,这些内容对内容生态系统的影响仍不明确。视频信息检索仍是获取视频的核心方式。基于检索模型在即时和图像检索任务中常偏好AI生成内容的现象,本文探究在更具挑战性的视频检索场景中是否也存在类似偏见,此时时空因素可能进一步影响模型行为。为此,我们构建了一个包含1.3万条由两种先进开源视频生成模型生成的视频的综合性基准数据集,并设计了一套公平严谨的评估指标以衡量偏见。该指标充分考虑了生成视频帧率有限与画质不佳可能带来的偏差。随后,我们使用三种现成的视频检索模型在此混合数据集上进行检索任务。结果表明,检索模型对生成视频存在明显偏好。进一步分析显示,将生成视频纳入检索模型的训练集会加剧这一偏见。与图像模态中的偏好不同,视频检索偏见源自未见过的视觉与时间信息的共同作用,其根源是两者的复杂交互。为缓解此偏见,我们采用对比学习对检索模型进行微调。本研究揭示了生成视频对检索系统潜在的重大影响。
原文摘要 · Abstract (English)
With the rapid development of AI-generated content (AIGC), the creation of high-quality AI-generated videos has become faster and easier, resulting in the Internet being flooded with all kinds of video content. However, the impact of these videos on the content ecosystem remains largely unexplored. Video information retrieval remains a fundamental approach for accessing video content. Building on the observation that retrieval models often favor AI-generated content in ad-hoc and image retrieval tasks, we investigate whether similar biases emerge in the context of challenging video retrieval, where temporal and visual factors may further influence model behavior. To explore this, we first construct a comprehensive benchmark dataset containing both real and AI-generated videos, along with a set of fair and rigorous metrics to assess bias. This benchmark consists of 13,000 videos generated by two state-of-the-art open-source video generation models. We meticulously design a suite of rigorous metrics to accurately measure this preference, accounting for potential biases arising from the limited frame rate and suboptimal quality of AIGC videos. We then applied three off-the-shelf video retrieval models to perform retrieval tasks on this hybrid dataset. Our findings reveal a clear preference for AI-generated videos in retrieval. Further investigation shows that incorporating AI-generated videos into the training set of retrieval models exacerbates this bias. Unlike the preference observed in image modalities, we find that video retrieval bias arises from both unseen visual and temporal information, making the root causes of video bias a complex interplay of these two factors. To mitigate this bias, we fine-tune the retrieval models using a contrastive learning approach. The results of this study highlight the potential implications of AI-generated videos on retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。