arXiv:2506.11727cs.IRcs.HC2025-06被引 5

YouTube搜索接口对学术研究不友好,数据不全且结果不稳定。

Forgetful by Design? A Critical Audit of YouTube's Search API for Academic Research

  • 用11个查询持续6个月测试,发现搜索结果不完整、不一致。
  • 发布后20-60天内视频几乎无法被检索到,时效性严重衰减。
  • 适合关注平台算法偏差或数据可信度的研究者参考。

本文系统审计了YouTube Data API v3的搜索接口,该接口常被用于学术研究。通过为期六个月的每周搜索,使用11个查询语句,我们发现其在完整性、代表性、一致性及偏见方面存在显著问题。结果显示,相关性与时间排序参数在视频召回率和准确率上差异巨大,相关性常返回大量无关视频。同时观察到严重的时效衰减:发布后20至60天内,可检索到的视频数量急剧下降,尽管这些视频仍在平台上存在。这可能破坏依赖系统性数据收集的研究设计。此外,相同查询在不同时间返回不同结果,导致可复现性受损。以欧洲议会选举为例的案例研究显示,这些问题会直接影响研究结论。尽管提出若干缓解策略,论文最终认为,该接口可能优先考虑‘新鲜度’而非全面检索,难以支持严谨的学术研究,尤其不符合《数字服务法案》要求。

原文摘要 · Abstract (English)

This paper critically audits the search endpoint of YouTube's Data API (v3), a common tool for academic research. Through systematic weekly searches over six months using eleven queries, we identify major limitations regarding completeness, representativeness, consistency, and bias. Our findings reveal substantial differences between ranking parameters like relevance and date in terms of video recall and precision, with relevance often retrieving numerous off-topic videos. We also observe severe temporal decay in video discoverability: the number of retrievable videos for a given period drops dramatically within just 20-60 days of publication, even though these videos remain on the platform. This potentially undermines research designs that rely on systematic data collection. Furthermore, search results lack consistency, with identical queries yielding different video sets over time, compromising replicability. A case study on the European Parliament elections highlights how these issues impact research outcomes. While the paper offers several mitigation strategies, it concludes that the API's search function, potentially prioritizing 'freshness' over comprehensive retrieval, is not adequate for robust academic research, especially concerning Digital Services Act requirements.

视频搜索数据质量学术研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。