arXiv:2506.18201cs.CLcs.CV2025-06被引 3

对比大模型在阿拉伯儿童绘本情绪识别中的表现

Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

  • 用三种提示策略测试GPT-4o和Gemini在阿拉伯绘本图像上的情绪识别
  • GPT-4o最高得分为59%宏F1,优于Gemini的43%
  • 模型常误判情绪极性,需加强文化敏感训练

多模态AI系统的情绪识别能力对开发文化响应型教育技术至关重要,但在阿拉伯语语境中仍研究不足,而此类适配学习工具亟需。本研究评估了GPT-4o与Gemini 1.5 Pro在处理阿拉伯儿童绘本插图时的情绪识别性能。基于普拉奇克情绪框架,使用7本阿拉伯故事书的75张图像,比较两种模型在零样本、少样本及思维链提示策略下的表现。GPT-4o在所有条件下均优于Gemini,思维链提示下达到最高宏F1-score为59%,而Gemini最佳为43%。错误分析显示,60.7%的错误源于情绪极性颠倒,且两者均难以处理文化特异性情绪与模糊叙事情境。研究揭示当前模型在文化理解上的根本局限,强调需采用文化敏感训练以开发面向阿拉伯语学习者的有效情绪感知教育技术。

原文摘要 · Abstract (English)

Emotion recognition capabilities in multimodal AI systems are crucial for developing culturally responsive educational technologies, yet remain underexplored for Arabic language contexts where culturally appropriate learning tools are critically needed. This study evaluates the emotion recognition performance of two advanced multimodal large language models, GPT-4o and Gemini 1.5 Pro, when processing Arabic children's storybook illustrations. We assessed both models across three prompting strategies (zero-shot, few-shot, and chain-of-thought) using 75 images from seven Arabic storybooks, comparing model predictions with human annotations based on Plutchik's emotional framework. GPT-4o consistently outperformed Gemini across all conditions, achieving the highest macro F1-score of 59% with chain-of-thought prompting compared to Gemini's best performance of 43%. Error analysis revealed systematic misclassification patterns, with valence inversions accounting for 60.7% of errors, while both models struggled with culturally nuanced emotions and ambiguous narrative contexts. These findings highlight fundamental limitations in current models' cultural understanding and emphasize the need for culturally sensitive training approaches to develop effective emotion-aware educational technologies for Arabic-speaking learners.

情绪识别多模态教育AI阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。