研究发现YouTube搜索接口结果波动大,难获取代表性历史数据。
On YouTube Search API Use in Research
- 12周内重复查询,发现搜索结果高度不稳定。
- 结果受话题热度影响,非高峰时段数据代表性差。
- 可能优先返回短且热门视频,适合关注趋势的研究者参考。
YouTube是全球使用最广泛的平台之一,近年来受到大量学术关注。尽管如此,人们对YouTube Data API尤其是搜索接口的实际运行机制仍了解有限。本文通过在12周内对同一查询进行多次测试,分析了该接口的行为。结果显示,搜索接口返回的结果在不同查询间存在显著差异。具体而言,API似乎会根据查询期间话题的相对热度对结果进行随机化处理,导致在非高峰话题时期几乎无法获得具有代表性的历史视频样本。此外,结果表明接口可能更倾向于返回较短且更受欢迎的视频,但频道热度的影响尚不明确。文章最后为研究人员提供了使用API进行数据采集的建议,并指出了未来扩展接口应用方向的研究路径。
原文摘要 · Abstract (English)
YouTube is among the most widely-used platforms worldwide, and has seen a lot of recent academic attention. Despite its popularity and the number of studies conducted on it, much less is understood about the way in which YouTube's Data API, and especially the Search endpoint, operates. In this paper, we analyze the API's behavior by running identical queries across a period of 12 weeks. Our findings show that the search endpoint returns highly variable results between queries. Specifically, the API seems to randomize returned videos based on the relative popularity of the respective topic during the query period, making it nearly impossible to obtain representative historical video samples, especially during non-peak topical periods. Our results also suggest that the API may prioritize shorter, more popular videos, although the role of channel popularity is not as clear. We conclude with suggested strategies for researchers using the API for data collection, as well as future research directions on expanding the API's use-cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。