提出文化忠实度评估框架,解决视频生成中文化准确性缺失问题
CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

- 将文化忠实度分解为身份、背景、行为三维度进行量化评估
- 最先进模型最高仅达56.8%得分,行为维度低于52.1%
- 人类偏好与视觉质量正相反,凸显文化准确性的关键性
随着Veo 3.1和LTX-2等视频生成模型的发展,其对全球多元文化的准确呈现仍是一个关键但研究不足的领域。现有指标如VideoScore仅衡量视觉质量,无法评估文化忠实度,导致错误手势(如用握手替代Namaste)与正确生成获得相同评分。本文提出CultureScore,一个分层评估框架,将文化忠实度拆解为身份(谁被呈现)、背景(文化特定环境)和行为(规范性动作与互动)三个细粒度维度。通过覆盖10个国家的评估套件,对三种顶尖模型生成的6,174个视频进行评测。结果表明:当前无一模型实现文化忠实生成,最佳模型整体得分仅56.8%,其中行为维度始终低于52.1%。此外,人工偏好排序与CultureScore方向一致,却与VideoScore相反——视觉质量最高的模型在人类评价中垫底,证明文化忠实度是实现公平视频生成的核心标准。数据与代码已公开于https://huggingface.co/datasets/ankurani/CultureScore。
原文摘要 · Abstract (English)
As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that replaces a Namaste with a handshake receives the same score as one that generates the gesture correctly. We propose CultureScore, a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions: Identity (who is represented), Context (culturally localized background), and Behavior (normative gestures and interactions). We operationalize this framework through an evaluation suite spanning 10 countries, yielding 6,174 generated videos across three state-of-the-art models. Our evaluation reveals that no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8\% overall CultureScore, with Behavior the most challenging dimension, which remains below 52.1\% across all models. Furthermore, human preference rankings align directionally with CultureScore but are inverted relative to VideoScore; the highest-scoring model on visual quality was ranked last by annotators, underscoring that cultural faithfulness is an essential criterion for equitable video generation. Data and code are available at https://huggingface.co/datasets/ankurani/CultureScore.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。