arXiv:2508.06388cs.CL2025-08被引 1

对比大模型与动漫迷,测试谁更会扮演角色并提供情感支持。

LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing

  • 构建首个情感支持型角色扮演数据集ChatAnime,含20位动漫角色与60个情绪场景。
  • 顶尖大模型在角色扮演和情感支持上超越人类,但人类在回应多样性上占优。
  • 适合研究虚拟角色、情感计算与人机交互的学者参考。

大型语言模型(LLMs)在角色扮演对话和情感支持方面已展现出色能力,但将两者结合以实现虚拟角色的情感支持互动仍存在研究空白。为此,本文以动漫角色为案例研究,因其性格鲜明且拥有庞大粉丝群体,便于评估模型在保持角色特质的同时提供情感支持的能力。我们提出ChatAnime,首个情感支持型角色扮演(ESRP)数据集:精心选取20位热门动漫角色,设计60个以情绪为中心的真实场景问题;通过全国选拔,招募40名精通特定角色的中国动漫爱好者,参与角色扮演。系统收集10个主流大模型与这40名爱好者两轮对话数据。为评估模型表现,设计涵盖基础对话、角色扮演、情感支持三个维度共9项细粒度指标,以及整体响应多样性指标。最终数据集包含2,400条人工撰写答案与24,000条大模型生成答案,附带超过13.2万条人工标注。实验表明,顶级大模型在角色扮演与情感支持上优于人类,而人类在响应多样性方面仍领先。本工作为优化大模型在情感支持角色扮演中的表现提供了宝贵资源与洞见。数据集已开源:https://github.com/LanlanQiu/ChatAnime。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive capabilities in role-playing conversations and providing emotional support as separate research directions. However, there remains a significant research gap in combining these capabilities to enable emotionally supportive interactions with virtual characters. To address this research gap, we focus on anime characters as a case study because of their well-defined personalities and large fan bases. This choice enables us to effectively evaluate how well LLMs can provide emotional support while maintaining specific character traits. We introduce ChatAnime, the first Emotionally Supportive Role-Playing (ESRP) dataset. We first thoughtfully select 20 top-tier characters from popular anime communities and design 60 emotion-centric real-world scenario questions. Then, we execute a nationwide selection process to identify 40 Chinese anime enthusiasts with profound knowledge of specific characters and extensive experience in role-playing. Next, we systematically collect two rounds of dialogue data from 10 LLMs and these 40 Chinese anime enthusiasts. To evaluate the ESRP performance of LLMs, we design a user experience-oriented evaluation system featuring 9 fine-grained metrics across three dimensions: basic dialogue, role-playing and emotional support, along with an overall metric for response diversity. In total, the dataset comprises 2,400 human-written and 24,000 LLM-generated answers, supported by over 132,000 human annotations. Experimental results show that top-performing LLMs surpass human fans in role-playing and emotional support, while humans still lead in response diversity. We hope this work can provide valuable resources and insights for future research on optimizing LLMs in ESRP. Our datasets are available at https://github.com/LanlanQiu/ChatAnime.

角色扮演情感支持大模型评测动漫数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。