arXiv:2601.18698cs.CV2026-01

评估视频生成模型对全球景点的视觉知识公平性,发现其表现比预期更均衡。

Are Video Generation Models Geographically Fair? An Attraction-Centric Evaluation of Global Visual Knowledge

  • 设计新框架GAP,从景点中心视角评估模型生成质量
  • 在500个全球景点上测试,模型对不同地区知识掌握较均匀
  • 适合关注AI地理公平性的研究者与应用开发者

近期文本到视频生成技术取得了令人惊叹的视觉效果,但这些模型是否具备地理上公平的视觉知识仍不明确。本文通过景点中心评估方法,探究文本到视频模型的地理公平性与地域化视觉知识。我们提出Geo-Attraction Landmark Probing(GAP)框架,系统评估模型对全球各地旅游景点的还原能力,并构建了包含500个分布广泛的景点的基准数据集GEOATTRACTION-500,覆盖不同区域和受欢迎程度。GAP融合多种互补指标,分离整体视频质量与景点特定知识,包括全局结构对齐、细粒度关键点对齐及视觉语言模型判断,均经人工评估验证。将GAP应用于当前最先进的Sora 2模型,结果表明:尽管普遍认为存在强地理偏差,但该模型在不同地区、发展水平和文化群体间表现出相对一致的地域化视觉知识,仅对景点热度有微弱依赖。这说明当前文本到视频模型在全球视觉知识表达上比预期更均衡,既展现了其在跨区域应用中的潜力,也提示需持续评估系统演进中的公平性。

原文摘要 · Abstract (English)

Recent advances in text-to-video generation have produced visually compelling results, yet it remains unclear whether these models encode geographically equitable visual knowledge. In this work, we investigate the geo-equity and geographically grounded visual knowledge of text-to-video models through an attraction-centric evaluation. We introduce Geo-Attraction Landmark Probing (GAP), a systematic framework for assessing how faithfully models synthesize tourist attractions from diverse regions, and construct GEOATTRACTION-500, a benchmark of 500 globally distributed attractions spanning varied regions and popularity levels. GAP integrates complementary metrics that disentangle overall video quality from attraction-specific knowledge, including global structural alignment, fine-grained keypoint-based alignment, and vision-language model judgments, all validated against human evaluation. Applying GAP to the state-of-the-art text-to-video model Sora 2, we find that, contrary to common assumptions of strong geographic bias, the model exhibits a relatively uniform level of geographically grounded visual knowledge across regions, development levels, and cultural groupings, with only weak dependence on attraction popularity. These results suggest that current text-to-video models express global visual knowledge more evenly than expected, highlighting both their promise for globally deployed applications and the need for continued evaluation as such systems evolve.

视频生成地理公平多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。