arXiv:2509.07403cs.CL2025-09被引 8

评测大模型在长文本中持续理解情绪的能力,发现新框架可显著提升表现。

LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction

  • 构建长上下文情绪评估基准,平均上下文达15,341词元
  • 引入多智能体协作与检索增强生成,提升情绪理解能力
  • 适合研究长对话、情感计算或模型鲁棒性的研究人员

大语言模型在情绪智能(EI)和长上下文建模方面取得显著进展,但现有评估基准常忽略情绪信息是随时间持续演化的长上下文过程。为填补长上下文推理中多维度情绪智能评估的空白,我们提出LongEmotion基准,涵盖情绪识别、知识应用与共情生成三类任务,平均上下文长度达15,341个词元。为提升真实场景下的性能,我们设计了协同情绪建模(CoEM)框架,结合检索增强生成(RAG)与多智能体协作,优化模型在长上下文中的情绪处理能力。我们对多种模型在长上下文设置下进行深入分析,考察推理模式激活、RAG检索策略及上下文长度适应性对其情绪智能表现的影响。

原文摘要 · Abstract (English)

Large language models (LLMs) have made significant progress in Emotional Intelligence (EI) and long-context modeling. However, existing benchmarks often overlook the fact that emotional information processing unfolds as a continuous long-context process. To address the absence of multidimensional EI evaluation in long-context inference and explore model performance under more challenging conditions, we present LongEmotion, a benchmark that encompasses a diverse suite of tasks targeting the assessment of models' capabilities in Emotion Recognition, Knowledge Application, and Empathetic Generation, with an average context length of 15,341 tokens. To enhance performance under realistic constraints, we introduce the Collaborative Emotional Modeling (CoEM) framework, which integrates Retrieval-Augmented Generation (RAG) and multi-agent collaboration to improve models' EI in long-context scenarios. We conduct a detailed analysis of various models in long-context settings, investigating how reasoning mode activation, RAG-based retrieval strategies, and context-length adaptability influence their EI performance. Our project page is: https://longemotion.github.io/

情绪智能长上下文多智能体RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。