arXiv:2409.01584cs.CL2024-09NAACL被引 6

跨语言艺术作品解释能力弱,英语知识难以迁移至其他语言。

Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models

  • 构建非机器翻译的多语言艺术解释数据集
  • 模型在非英语语言上表现显著下降
  • 英语指令微调无法有效提升其他语言性能

随着大规模视觉语言模型(LVLMs)性能提升,其多语言生成能力备受期待。然而,视觉编码器预训练及与语言模型联合训练主要依赖英语数据,导致模型在非英语语言上的解释能力尚不明确。现有跨语言问答基准多通过机器翻译构建,存在文化偏差。为此,本研究构建了无需机器翻译的多语言艺术解释数据集,包含文化细节与地域表达。实验表明,LVLMs在非英语语言上表现明显劣于英语;且模型难以有效迁移英语中学到的知识。数据集已公开:https://huggingface.co/datasets/naist-nlp/MultiExpArt。

原文摘要 · Abstract (English)

As the performance of Large-scale Vision Language Models (LVLMs) improves, they are increasingly capable of responding in multiple languages, and there is an expectation that the demand for explanations generated by LVLMs will grow. However, pre-training of Vision Encoder and the integrated training of LLMs with Vision Encoder are mainly conducted using English training data, leaving it uncertain whether LVLMs can completely handle their potential when generating explanations in languages other than English. In addition, multilingual QA benchmarks that create datasets using machine translation have cultural differences and biases, remaining issues for use as evaluation tasks. To address these challenges, this study created an extended dataset in multiple languages without relying on machine translation. This dataset that takes into account nuances and country-specific phrases was then used to evaluate the generation explanation abilities of LVLMs. Furthermore, this study examined whether Instruction-Tuning in resource-rich English improves performance in other languages. Our findings indicate that LVLMs perform worse in languages other than English compared to English. In addition, it was observed that LVLMs struggle to effectively manage the knowledge learned from English data. Our dataset is available at https://huggingface.co/datasets/naist-nlp/MultiExpArt

跨语言视觉语言模型艺术解释多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。