评测视频大模型跨文化理解能力,发现中文场景表现更差。
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
- 用中美电视剧构建标注数据集,结合AI初筛与人工复核
- 模型在中文文化理解上表现显著低于美国文化,非语言线索更难捕捉
- 视频模态不可替代,模型越大效果不一定更好
随着视频大语言模型(VideoLLMs)在全球部署,评估其跨文化推理能力至关重要。为此,我们提出VideoNorms数据集,基于流行的美中电视节目,对文化规范的遵守或违背进行标注,并提供(非)言语证据。通过人机协作框架,每个样本先由大模型初标,再经至少三位具有目标文化生活经验的单文化标注员审核,最终获得超过3000条人工判断。人类验证显示,中美文化规范提取性能存在差异,警示训练数据中代表性不足的文化不应依赖全自动方法。对7个开源视频大模型的分层线性建模分析表明:1)模型在中文文化中的表现劣于美国文化,尤其在规范遵守预测上;2)模型在提供非言语证据方面比言语证据更困难。消融实验确认视频模态对准确性能至关重要,且扩大模型规模未能提升分类得分。研究结果与数据为更具文化根基的视频模型训练与评估提供支持。
原文摘要 · Abstract (English)
As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNorms, a dataset of cultural norm annotations from popular US and Chinese TV shows annotated with adherence or violation labels and (non-)verbal evidence. Through a human-AI collaboration framework, each item was first annotated by a large VideoLLM, and then reviewed by at least three trained monocultural annotators with significant lived experience in the target culture, resulting in a dataset of over 3,000 human judgments. Human verification showed disparity in US and Chinese norm extraction performance, cautioning against fully automatic approaches cultures under-represented in training data. Hierarchical linear modeling analysis of $7$ open-weight VideoLLMs' performance revealed that: 1) models perform worse in Chinese compared to US, particularly for norm adherence prediction; 2) models have more difficulty in providing non-verbal evidence compared to verbal evidence for norm adherence/violation predictions. Ablation studies confirm video modality is indeed necessary for accurate performance, and scaling model size does not yield classification score improvements. Our findings and data contribute to culturally grounded video model training and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。