用三个协作智能体提升大模型对卡通图像的问答理解能力
Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA
- 设计视觉、语言、批判三类智能体协同推理
- 在Pororo和Simpsons数据集上验证框架有效性
- 揭示多智能体在卡通图文理解中的行为机制
针对风格化卡通图像的视觉问答(VQA)任务,标准大语言模型因训练数据以自然图像为主,难以应对夸张的视觉抽象与叙事性上下文。为此,本文提出一种专为卡通图像VQA设计的多智能体大语言模型框架,包含视觉、语言与批判三类专用智能体,通过协同工作整合视觉线索与叙事背景,实现结构化推理。该框架在两个基于卡通的VQA数据集——Pororo与Simpsons上进行系统评估,实验结果详细分析了各智能体对最终预测的贡献,深入揭示了大模型在卡通图像理解与多模态推理中多智能体行为的运作机制。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) for stylised cartoon imagery presents challenges, such as interpreting exaggerated visual abstraction and narrative-driven context, which are not adequately addressed by standard large language models (LLMs) trained on natural images. To investigate this issue, a multi-agent LLM framework is introduced, specifically designed for VQA tasks in cartoon imagery. The proposed architecture consists of three specialised agents: visual agent, language agent and critic agent, which work collaboratively to support structured reasoning by integrating visual cues and narrative context. The framework was systematically evaluated on two cartoon-based VQA datasets: Pororo and Simpsons. Experimental results provide a detailed analysis of how each agent contributes to the final prediction, offering a deeper understanding of LLM-based multi-agent behaviour in cartoon VQA and multimodal inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。