arXiv:2502.13012cs.HCcs.CL2025-02ACL综述被引 30

提出一套可落地的RPA评估指南,解决智能体评价标准不统一问题。

Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents

  • 系统分析1676篇论文,提炼出六类智能体属性
  • 归纳七类任务特征与七种评估指标,构建完整评估框架
  • 适合研究RPA的开发者与评测人员参考使用

角色扮演智能体(RPA)是近年来基于大语言模型的热门研究方向,能模拟多样任务中的人类行为。然而,由于任务需求和智能体设计的多样性,对RPA的评估面临挑战。本文通过系统回顾2021年1月至2024年12月间发表的1,676篇论文,提出一套基于证据、可操作且具普适性的大语言模型驱动型RPA评估设计指南。分析识别出六类智能体属性、七类任务属性及七种评估指标,据此构建了结构化评估框架,助力研究者建立更系统、一致的评估方法。

原文摘要 · Abstract (English)

Role-Playing Agent (RPA) is an increasingly popular type of LLM Agent that simulates human-like behaviors in a variety of tasks. However, evaluating RPAs is challenging due to diverse task requirements and agent designs. This paper proposes an evidence-based, actionable, and generalizable evaluation design guideline for LLM-based RPA by systematically reviewing 1,676 papers published between Jan. 2021 and Dec. 2024. Our analysis identifies six agent attributes, seven task attributes, and seven evaluation metrics from existing literature. Based on these findings, we present an RPA evaluation design guideline to help researchers develop more systematic and consistent evaluation methods.

角色扮演智能体评估LLM综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。