为AI情绪智能评估构建新框架,聚焦可衡量能力而非人类体验。
Why We Need a New Framework for Emotional Intelligence in AI
- 区分人类情绪智能中不可复制的主观体验与可评估的感知响应能力。
- 指出现有评测基准缺乏对情绪本质的理论支撑,导致评估不全面。
- 适合关注AI情感交互、伦理评估与评测标准的研究者参考。
本文主张当前人工智能(AI)情绪智能(EI)评估框架亟需改进,因其未能充分涵盖与AI相关的EI各个维度。人类的情绪智能常包含现象学层面的体验和理解感,而这些是人工系统所不具备的,因此部分方面不适用于评估AI。然而,情绪智能也包含感知情绪状态、解释情绪、适当回应及适应新情境(如跨文化环境)的能力,这些在不同程度上是人工系统可以实现的。尽管已有若干基准框架专注于评估不同AI模型在相关任务中的表现,但它们普遍缺乏对情绪本质及情绪智能内涵的扎实理论基础。本研究首先回顾了情绪与一般情绪智能的不同理论,评估其在人工智能系统中的适用性;随后批判性分析现有基准框架,揭示其在第一部分提出的EI理论框架下的不足;最后,提出若干改进评估策略的选项,以克服当前情绪智能评估中的缺陷。
原文摘要 · Abstract (English)
In this paper, we develop the position that current frameworks for evaluating emotional intelligence (EI) in artificial intelligence (AI) systems need refinement because they do not adequately or comprehensively measure the various aspects of EI relevant in AI. Human EI often involves a phenomenological component and a sense of understanding that artificially intelligent systems lack; therefore, some aspects of EI are irrelevant in evaluating AI systems. However, EI also includes an ability to sense an emotional state, explain it, respond appropriately, and adapt to new contexts (e.g., multicultural), and artificially intelligent systems can do such things to greater or lesser degrees. Several benchmark frameworks specialize in evaluating the capacity of different AI models to perform some tasks related to EI, but these often lack a solid foundation regarding the nature of emotion and what it is to be emotionally intelligent. In this project, we begin by reviewing different theories about emotion and general EI, evaluating the extent to which each is applicable to artificial systems. We then critically evaluate the available benchmark frameworks, identifying where each falls short in light of the account of EI developed in the first section. Lastly, we outline some options for improving evaluation strategies to avoid these shortcomings in EI evaluation in AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。