RPA-Check可自动评估角色扮演大模型的稳定性与合规性,解决传统指标无法衡量的问题。
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents

- 构建四阶段框架,将角色行为转化为可量化检查项
- 在5个法律场景中发现小模型(8-9B)比大模型更稳定可靠
- 适合需要高合规性的角色扮演系统开发者与评测人员
大型语言模型在交互系统中的快速应用催生了动态、开放的角色扮演代理(RPAs)。然而,评估这些代理仍面临挑战,因标准NLP指标难以捕捉角色一致性、逻辑连贯性及长期叙事稳定性。本文提出RPA-Check,一种多阶段自动化评估框架,用于客观评测复杂约束环境下的LLM-RPAs表现。方法包括:(1) 定义维度,建立高层次行为标准;(2) 增强,将要求细化为细粒度布尔检查项;(3) 语义过滤,确保指标客观、无冗余且代理隔离;(4) 利用链式思维验证的LLM-as-a-Judge评分机制。通过在包含多个量化本地模型的法律训练游戏LLM Court上验证,五种不同法律场景的实验结果表明,该框架能识别出模型规模、推理深度与运行稳定性之间的微妙权衡。值得注意的是,参数量与程序一致性呈反比关系,较小但充分指令微调的模型(8-9B)反而优于易受用户对齐偏差或阿谀倾向影响的大架构。RPA-Check为特定领域生成式代理评估提供了标准化、可复现的度量基准。
原文摘要 · Abstract (English)
The rapid adoption of Large Language Models (LLMs) in interactive systems has enabled the creation of dynamic, open-ended Role-Playing Agents (RPAs). However, evaluating these agents remains a significant challenge, as standard NLP metrics fail to capture the nuances of role adherence, logical consistency, and long-term narrative stability. This paper introduces RPA-Check, a multi-stage automated evaluation framework designed to objectively assess the performance of LLM-based RPAs in complex, constraints-heavy environments. Our methodology is based on a four-step pipeline: (1) Dimension Definition, establishing high-level qualitative behavioral criteria; (2) Augmentation, where these requirements are expanded into granular boolean checklist indicators; (3) Semantic Filtering, to ensure indicator objectivity, no redundancy and agent isolation; and (4) LLM-as-a-Judge Evaluation, which employs chain-of-thought verification to score agent fidelity. We validate this framework by applying it to LLM Court, a serious game for forensic training involving several quantized local models. Experimental results across five distinct legal scenarios demonstrate the framework's ability to identify subtle trade-offs between model size, reasoning depth, and operational stability. Notably, the findings reveal an inverse relationship between parametric scale and procedural consistency, showing that smaller, adequately instruction-tuned models (8-9B) can outperform larger architectures prone to user-alignment bias or sycophancy. RPA-Check thus provides a standardized and reproducible metric for future research in generative agent evaluation within specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。