arXiv:2506.01565cs.CLcs.CV2025-06EMNLP被引 4

构建汉服跨时文化理解与转译的多模态评测基准

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

  • 以汉服为载体,设计跨时文化理解与现代转译任务
  • 闭源模型比人类非专家差10%,最佳模型转译成功率仅42%
  • 适合研究文化理解、跨时代视觉生成与文化遗产数字化的学者

文化是随地理与时间演化的动态领域。现有视觉语言模型对文化理解的研究多聚焦地理多样性,忽视了时间维度。为此,我们提出Hanfu-Bench,一个由专家精心构建的多模态基准数据集。汉服作为贯穿中国古代王朝的传统服饰,既体现中华文化的深层时间性,又在当代中国社会广受欢迎。该基准包含两大任务:文化视觉理解与文化图像转译。前者通过单图或多图选择题问答评估时间文化特征识别能力;后者关注传统服饰向现代设计的转化,需兼顾文化元素传承与现代语境适配。评估显示,闭源模型在视觉文化理解上表现接近非专家,但较人类专家低10%;开源模型更差,甚至低于非专家。在转译任务中,多维度人工评估表明最优模型成功率仅为42%。该基准揭示了跨时文化理解与创造性适配中的重大挑战,为该方向研究提供关键测试平台。

原文摘要 · Abstract (English)

Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions. To bridge this gap, we introduce Hanfu-Bench, a novel, expert-curated multimodal dataset. Hanfu, a traditional garment spanning ancient Chinese dynasties, serves as a representative cultural heritage that reflects the profound temporal aspects of Chinese culture while remaining highly popular in Chinese contemporary society. Hanfu-Bench comprises two core tasks: cultural visual understanding and cultural image transcreation. The former task examines temporal-cultural feature recognition based on single- or multi-image inputs through multiple-choice visual question answering, while the latter focuses on transforming traditional attire into modern designs through cultural element inheritance and modern context adaptation. Our evaluation shows that closed VLMs perform comparably to non-experts on visual cutural understanding but fall short by 10% to human experts, while open VLMs lags further behind non-experts. For the transcreation task, multi-faceted human evaluation indicates that the best-performing model achieves a success rate of only 42%. Our benchmark provides an essential testbed, revealing significant challenges in this new direction of temporal cultural understanding and creative adaptation.

文化理解多模态汉服转译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。