多智能体协作提升跨文化图像描述准确性
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
- 设计多智能体框架,每个智能体代表不同文化背景
- 在中、印、罗三地图像上实现更准确的文化相关描述
- 提出可衡量文化信息的评估指标,适合跨文化研究者
大型多模态模型(LMMs)在多种多模态任务中表现优异,但在跨文化场景下受限于数据与模型的西方中心倾向。本文探索了多智能体协同在跨文化图像描述任务中的潜力。提出MosAIC框架,通过具有不同文化身份的LMM智能体进行协作;构建涵盖中国、印度、罗马尼亚三地的英文文化增强图像标注数据集,覆盖GeoDE、GD-VCR、CVQA三个数据集;提出一种可适应文化的评估指标,用于衡量描述中文化信息的丰富性;实验表明,多智能体交互在各项指标上均优于单智能体模型,为未来研究提供重要启示。代码与数据已开源。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) exhibit impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of most data and models. Conversely, multi-agent models have shown significant capability in solving complex tasks. Our study evaluates the collective performance of LMMs in a multi-agent interaction setting for the novel task of cultural image captioning. Our contributions are as follows: (1) We introduce MosAIC, a Multi-Agent framework to enhance cross-cultural Image Captioning using LMMs with distinct cultural personas; (2) We provide a dataset of culturally enriched image captions in English for images from China, India, and Romania across three datasets: GeoDE, GD-VCR, CVQA; (3) We propose a culture-adaptable metric for evaluating cultural information within image captions; and (4) We show that the multi-agent interaction outperforms single-agent models across different metrics, and offer valuable insights for future research. Our dataset and models can be accessed at https://github.com/MichiganNLP/MosAIC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。