arXiv:2608.01942cs.CVcs.CL2026-08被引 2

评测视频生成模型对文化细节的理解能力,发现主流模型仍存在文化失真。

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

论文配图:CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
图 1 · 摘自论文原文
  • 构建涵盖12国8区域的文化多样性视频评测集
  • 模型在文化细节还原上表现不佳,尤其对小众文化
  • 强调动态多模态文化表达,适合跨文化研究者使用

文本到视频(T2V)生成模型发展迅速,但其对多元文化语境的呈现能力仍缺乏系统评估。现有基准主要关注感知质量、物理合理性与文本-视频对齐,却未直接考察生成视频是否准确呈现文化特有物体、行为、仪式、可见文字或音频线索。本文提出CultureVidBench,一个全面评估T2V模型文化理解能力的基准。该数据集包含1000个精心筛选的提示,覆盖12个国家、6大洲、8个文化区域和14个文化维度,分为物质文化、社会实践与表演、仪式与典礼三类。针对视频生成特性,强调动态与多模态文化表征,包括社会互动、仪式流程及符合文化的可见文字与音频。我们通过人工用户研究与基于多模态大语言模型(MLLM)的自动评估,对七种代表性T2V模型进行文化忠实度、多模态文化呈现、语义一致性与感知质量的评测。结果表明,尽管当前模型在语义一致性和视觉质量上表现良好,但在捕捉细微文化细节方面普遍不足,尤其在非主流地区、仪式场景及多模态文化线索上表现更差。

原文摘要 · Abstract (English)

Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.

视频生成文化理解多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。