评估AI生成故事在非洲五社区的文化契合度,发现语境比符号更重要
Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities

- 基于五个非洲社区的混合方法评估,结合定量标注与定性讨论
- 文化契合度依赖社会、语言、流程和视觉语境,非仅靠可识别符号
- 提出文化对齐分类体系,建议本地化校准自动化评估流程
本文研究AI生成的多模态故事在五个非洲社区中的文化契合度,考察其是否符合当地的生活实践、人际关系、语言习惯、价值观念及视觉预期。通过19位文化代表参与的扎根社区混合方法评估,结合量化标注与质性焦点小组讨论,发现文化契合不仅取决于可识别的文化标记,更在于这些标记如何融入社会、语言、程序和视觉语境。基于评估结果,构建了包含五类文化标记与八种常见错配机制的文化对齐分类体系。同时评估了五个多模态大模型判别器在规模化评估中的表现,发现其可靠性与评分校准在不同社区间差异显著,无单一判别器在所有场景中表现一致。该研究推动建立以社区判断为基准的校准评估流程,明确自动化工具可信赖与需人工介入的边界。
原文摘要 · Abstract (English)
In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community-grounded judgments at scale. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings. These findings motivate community-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。