测试大模型能否像人一样识别角色对话,发现其准确率远低于人类。
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- 构建首个评测基准PersonaEval,检验模型对角色对话的识别能力。
- 顶尖大模型角色识别准确率仅69%,远低于人类的90.8%。
- 当前大模型缺乏人类级推理能力,难以胜任角色扮演评估任务。
当前角色扮演研究多依赖未经验证的大模型作为评判者,可能无法反映人类对角色一致性的感知。判断角色扮演质量的前提是正确识别说话者身份,即角色识别能力。本文提出PersonaEval,首个专门测试大模型角色识别能力的基准,使用小说、剧本和视频字幕中的人类创作对话,要求模型根据语境判断说话角色。实验包括人类对照研究,结果显示即使表现最佳的大模型准确率也仅达69%,远低于人类参与者的90.8%。这表明当前大模型尚不具备人类般的角色理解能力。进一步分析训练与推理阶段因素,发现可靠评估不仅需任务微调,更依赖强健的人类级推理能力。相关数据集已开源:https://github.com/maple-zhou/PersonaEval。
原文摘要 · Abstract (English)
Current role-play studies often rely on unvalidated LLM-as-a-judge paradigms, which may fail to reflect how humans perceive role fidelity. A key prerequisite for human-aligned evaluation is role identification, the ability to recognize who is speaking based on dialogue context. We argue that any meaningful judgment of role-playing quality (how well a character is played) fundamentally depends on first correctly attributing words and actions to the correct persona (who is speaking). We present PersonaEval, the first benchmark designed to test whether LLM evaluators can reliably identify human roles. PersonaEval uses human-authored dialogues from novels, scripts, and video transcripts, challenging models to determine the correct persona according to the conversation context. Our experiments, including a human study, show that even the best-performing LLMs reach only around 69% accuracy, well below the level needed for reliable evaluation. In contrast, human participants perform near ceiling with 90.8% accuracy, highlighting that current LLM evaluators are still not human enough to effectively judge role-play scenarios. To better understand this gap, we examine training-time adaptation and test-time compute, suggesting that reliable evaluation requires more than task-specific tuning, but depends on strong, human-like reasoning abilities in LLM evaluators. We release our benchmark at https://github.com/maple-zhou/PersonaEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。