首个评估具身智能体共情行为的基准,验证其能否像人一样回应情感需求。
EmpathyAgent: Can Embodied Agents Conduct Empathetic Actions?
- 构建包含1万组多模态数据的共情任务基准,支持跨场景评估。
- 现有模型在共情行为上表现有限,仅能完成部分情感响应任务。
- 基于该基准微调的Llama3-8B展现出提升共情能力的潜力,适合心理陪伴研究者使用。
共情是人类互动的核心,但具身智能体是否能提供类人共情支持仍不明确。现有研究关注智能体的任务解决与社交能力,却忽视了其理解共情需求并执行共情行为的能力。为此,我们提出EmpathyAgent,首个用于评估和提升智能体在多样情境中进行共情行为的基准。该基准包含10,000个多媒体样本,配有对应共情任务计划,并设置三种不同挑战。为系统评估共情行为,我们设计了一个专属评估套件,涵盖共情过程的多个维度。我们对当前主流模型进行基准测试,发现实现真正共情行为仍是重大挑战。同时,我们使用EmpathyAgent训练Llama3-8B模型,结果表明其具备增强共情行为的潜力。通过建立共情行为的标准化评估体系,我们期望推动具身共情智能体的研究进展。代码与数据已公开于https://github.com/xinyan-cxy/EmpathyAgent。
原文摘要 · Abstract (English)
Empathy is fundamental to human interactions, yet it remains unclear whether embodied agents can provide human-like empathetic support. Existing works have studied agents' tasks solving and social interactions abilities, but whether agents can understand empathetic needs and conduct empathetic behaviors remains overlooked. To address this, we introduce EmpathyAgent, the first benchmark to evaluate and enhance agents' empathetic actions across diverse scenarios. EmpathyAgent contains 10,000 multimodal samples with corresponding empathetic task plans and three different challenges. To systematically evaluate the agents' empathetic actions, we propose an empathy-specific evaluation suite that evaluates the agents' empathy process. We benchmark current models and found that exhibiting empathetic actions remains a significant challenge. Meanwhile, we train Llama3-8B using EmpathyAgent and find it can potentially enhance empathetic behavior. By establishing a standard benchmark for evaluating empathetic actions, we hope to advance research in empathetic embodied agents. Our code and data are publicly available at https://github.com/xinyan-cxy/EmpathyAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。