arXiv:2508.04199cs.CL2025-08被引 2

测试大模型在肯尼亚青年聊天中的情感推理能力,发现顶尖模型表现更稳定。

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

  • 将情感视为文化嵌入的动态概念,用反事实和人工标注数据评估模型
  • 顶级模型在模糊或情绪反转时仍保持推理一致性,开源模型则易出错
  • 适合关注跨文化AI评估与社会语义理解的研究者

低资源、文化复杂的语境下,传统自然语言处理方法因假设固定标签和普适情感表达而面临挑战。本文提出一种诊断框架,将情感视为依赖语境、文化嵌入的构造,并评估大型语言模型(LLMs)在内罗毕青年健康群组非正式、混合编码的WhatsApp消息中的情感推理能力。通过结合人工标注数据、情感反转的反事实样本及基于量规的解释评估,我们探查了模型的可解释性、鲁棒性与人类推理的一致性。以社会科学测量视角重构评估,将模型输出视为测量抽象情感概念的工具。研究发现,模型推理质量存在显著差异:顶级大模型表现出解释稳定性,而开放模型在歧义或情感转换下常失效。本工作强调了在复杂现实交流中需采用文化敏感且具备推理意识的AI评估范式。

原文摘要 · Abstract (English)

Sentiment analysis in low-resource, culturally nuanced contexts challenges conventional NLP approaches that assume fixed labels and universal affective expressions. We present a diagnostic framework that treats sentiment as a context-dependent, culturally embedded construct, and evaluate how large language models (LLMs) reason about sentiment in informal, code-mixed WhatsApp messages from Nairobi youth health groups. Using a combination of human-annotated data, sentiment-flipped counterfactuals, and rubric-based explanation evaluation, we probe LLM interpretability, robustness, and alignment with human reasoning. Framing our evaluation through a social-science measurement lens, we operationalize and interrogate LLMs outputs as an instrument for measuring the abstract concept of sentiment. Our findings reveal significant variation in model reasoning quality, with top-tier LLMs demonstrating interpretive stability, while open models often falter under ambiguity or sentiment shifts. This work highlights the need for culturally sensitive, reasoning-aware AI evaluation in complex, real-world communication.

情感分析文化敏感大模型评估低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。