评测多模态大模型跨语言一致性,发现主流模型仍存短板
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
- 设计双基准:知识回忆与视觉记忆,覆盖15种语言
- 跨语言知识准确率普遍低于60%,文化相关问题表现更差
- 适合关注多语言AI公平性与文化适配的研究者
多模态大语言模型(MLLMs)的快速发展显著提升了实际应用能力,但在跨语言情境下保持性能一致性,尤其是整合文化知识方面,仍是重大挑战。为更好评估该问题,我们引入两个新基准:KnowRecall 和 VisRecall,用于衡量 MLLMs 在跨语言场景下的表现。KnowRecall 是一个视觉问答基准,测试15种语言中关于全球地标的历史与文化事实知识一致性。VisRecall 则通过在无图条件下要求模型用9种语言描述地标外观,评估视觉记忆一致性。实验结果表明,包括专有模型在内的先进 MLLMs 在跨语言一致性上仍存在明显不足,凸显出构建真正多语言、文化感知型模型的迫切需求。
原文摘要 · Abstract (English)
The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural knowledge, remains a significant challenge. To better assess this issue, we introduce two new benchmarks: KnowRecall and VisRecall, which evaluate cross-lingual consistency in MLLMs. KnowRecall is a visual question answering benchmark designed to measure factual knowledge consistency in 15 languages, focusing on cultural and historical questions about global landmarks. VisRecall assesses visual memory consistency by asking models to describe landmark appearances in 9 languages without access to images. Experimental results reveal that state-of-the-art MLLMs, including proprietary ones, still struggle to achieve cross-lingual consistency. This underscores the need for more robust approaches that produce truly multilingual and culturally aware models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。